RAG
DeepResearch-9K: A Challenging Benchmark Dataset of Deep-Research Agent
The article announces the release of DeepResearch-9K, a challenging benchmark dataset designed for deep-research agents, featuring 9,000 multi-step questions across three difficulty levels, high-quality search trajectories from the Tongyi-DeepResearch-30B-A3B model, and verifiable answers. It also introduces the open-source framework DeepResearch-R1, which supports multi-turn web interactions and various reinforcement learning approaches. This release is significant for practitioners as it provides a robust dataset and framework to enhance the training and evaluation of deep-research agents, addressing the current limitations in available resources.
datasetdeep-researchqa