Safety
ComplexConstraints and Beyond: Expert Rubrics for RLVR
The article introduces ComplexConstraints, a new expert-curated instruction-following dataset designed to enhance evaluation methods for large language models (LLMs). It details five design principles for creating high-quality rubrics and demonstrates that training on approximately 1,000 examples from this dataset leads to significant performance improvements: +15.5% for a 4B-parameter model and +12.2% for a 235B-parameter model on instruction-following tasks. This approach not only improves evaluation accuracy but also serves as an effective reinforcement learning training signal, benefiting model performance on out-of-distribution benchmarks.
evaluationrubricsllm