#177 · Primary category: MLOps & Evaluation

can-ai-code

ai ggml humaneval langchain llama-cpp llm transformers

Self-evaluating interview for AI coders

Project last updated:06/21/25

GitHub Stars

598

Forks

35

Contributors

9

License

MIT

Why we included this project

Most coding benchmarks stop being useful once models get good enough to ace them, and this project was built to dodge that ceiling. Instead of fixed test cases, it generates unlimited unique problems and scores how far up a two-dimensional difficulty ramp each model can climb, where one axis adds working-memory load and the other adds structural complexity. That makes it a practical way to compare models side by side or track a single model's progress over time without everything clustering at the top. It also ships with results from many model families, so you can get a sense of where models stand before spending your own compute.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category