Open Source
CuratorKIT : Data Curation and Synthetic Data Generation for LLM Post-Training
CuratorKIT is an open-source Python library designed for data curation and synthetic data generation in post-training workflows for large language models. It integrates ingestion, deduplication, synthetic generation, and quality filtering into a single configurable pipeline, featuring LLM-powered generation tasks, provenance tracking, and structured failure reporting. This tool supports over 100 LLM providers and offers both a Python API and a YAML-driven CLI, making it essential for practitioners seeking reproducible and auditable data management at scale.
data curationllmpipeline