#105 · Primary category: AI Tool Directories & Curated Lists

Awesome-Code-LLM

ai awesome datasets llm nlp papers software-engineering survey tmlr

[TMLR] A curated list of language modeling researches for code (and other software engineering activities), plus related datasets.

Project last updated:05/20/26

GitHub Stars

3.4K

Forks

237

Contributors

15

License

Other

Why we included this project

This is a curated reading list for anyone working on language models for code, assembled by the team behind a peer-reviewed TMLR survey. Instead of a flat pile of links, it organizes hundreds of papers and datasets into a taxonomy spanning pretraining, fine-tuning, reinforcement learning, code agents, and downstream tasks like generation, repair, testing, and DevOps. That structure makes it easy for researchers and engineers to find the work relevant to a specific problem, whether that's code search, program repair, or text-to-SQL, without wading through unrelated material. The dataset and benchmark sections are a practical bonus for people planning experiments, since the evaluation resources are collected in one place. The list is actively maintained, with recent conference papers folded in, so it stays useful as the field moves quickly.

Articles for this project

No articles for this project yet.

To suggest a topic or contribute an article, contact us.

Related projects in this category