#177 · Primary category: MLOps & Evaluation
can-ai-code
Self-evaluating interview for AI coders
Project last updated:06/21/25
GitHub Stars
598
Forks
35
Contributors
9
License
MIT
Why we included this project
Most coding benchmarks stop being useful once models get good enough to ace them, and this project was built to dodge that ceiling. Instead of fixed test cases, it generates unlimited unique problems and scores how far up a two-dimensional difficulty ramp each model can climb, where one axis adds working-memory load and the other adds structural complexity. That makes it a practical way to compare models side by side or track a single model's progress over time without everything clustering at the top. It also ships with results from many model families, so you can get a sense of where models stand before spending your own compute.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
unsloth
Local UI to run and train LLMs and diffusion models. Supports GGUF, MLX, Qwen3.8, Kimi K3, MiniMax-H3, Gemma 4, FLUX and more.
LlamaFactory
Unified Efficient Fine-Tuning of 100+ LLMs & VLMs (ACL 2024)
airflow
Apache Airflow - A platform to programmatically author, schedule, and monitor workflows
langfuse
🪢 Open source AI engineering platform: LLM evals, observability, metrics, prompt management, playground, datasets. Integrates with OpenTelemetry, LangChain, OpenAI SDK, LiteLLM, and more. 🍊YC W23
netron
Visualizer for neural network, deep learning and machine learning models