RAG
DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
DoGMaTiQ is a newly introduced pipeline designed for the automated generation of question-and-answer nuggets to evaluate long-form, citation-backed reports, particularly in cross-lingual contexts. The pipeline consists of three stages: document-grounded nugget generation, paraphrase clustering, and nugget subselection based on quality criteria, and it integrates with the AutoArgue framework for automatic report evaluation. Extensive experiments on TREC shared tasks demonstrate strong rank correlations with human evaluations, highlighting the importance of a robust LLM nugget generator in the evaluation process, with the code and artifacts made publicly available for further research.
evaluationnuggetsreport-evaluation