#64 · Primary category: MLOps & Evaluation
Kiln
Build, Evaluate, and Optimize AI Systems. Includes evals, RAG, agents, fine-tuning, synthetic data generation, dataset management, MCP, and more.
Project last updated:08/30/26
GitHub Stars
5.0K
Forks
375
Contributors
14
License
Other
Why we included this project
Kiln pairs a desktop workbench with an MIT-licensed Python library, covering the AI development loop from evaluation to fine-tuning instead of just one stage. Define a task once and the same dataset flows through automated evaluation, prompt optimization, RAG, and synthetic data generation, and the auto-optimizer searches hundreds of prompt mutations and model choices rather than just reporting a score. Product managers and subject experts can rate outputs and flag regressions in the GUI while engineers ship the same task to production through the library, and the eval builder generates judge prompts and synthetic datasets in roughly ten minutes. It runs locally with your own API keys or fully offline through Ollama, and work syncs over the Git infrastructure you already use.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models