ai-digest.dev
last updated 4 h ago
Open SourcearXiv cs.CL 34 d ago

CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training

CuratorKIT is an open-source Python library designed for data curation and synthetic data generation in post-training workflows for large language models. It integrates ingestion, deduplication, synthetic generation, and quality filtering into a single configurable pipeline, featuring LLM-powered generation tasks, provenance tracking, and structured failure reporting. This tool supports over 100 LLM providers and offers both a Python API and a YAML-driven CLI, making it essential for practitioners seeking reproducible and auditable data management at scale.

data curationllmpipelinerelevance 0.00 · engagement 0.00
Read at source ↗← all news