#64 · Primary category: MLOps & Evaluation

Kiln

ai chain-of-thought collaboration dataset-generation evals evaluation evaluation-framework fine-tuning machine-learning macos mcp ml ollama openai prompt prompt-engineering python rlhf synthetic-data windows

Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.

Project last updated:08/30/26

GitHub Stars

5.0K

Forks

375

Contributors

14

License

Other

Why we included this project

Kiln pairs a desktop workbench with an MIT-licensed Python library, covering the AI development loop from evaluation to fine-tuning instead of just one stage. Define a task once and the same dataset flows through automated evaluation, prompt optimization, RAG, and synthetic data generation, and the auto-optimizer searches hundreds of prompt mutations and model choices rather than just reporting a score. Product managers and subject experts can rate outputs and flag regressions in the GUI while engineers ship the same task to production through the library, and the eval builder generates judge prompts and synthetic datasets in roughly ten minutes. It runs locally with your own API keys or fully offline through Ollama, and work syncs over the Git infrastructure you already use.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category