#84 · Primary category: Foundation Models
DeepSeek-LLM
DeepSeek LLM: Let there be answers
Project last updated:02/04/24
GitHub Stars
7.3K
Forks
1.3K
Contributors
10
License
MIT
Why we included this project
DeepSeek-LLM holds the 7B and 67B base and chat checkpoints DeepSeek trained from scratch on 2 trillion English and Chinese tokens, with download links and quick-start notes in the README. Teams that want a large bilingual model they can run and fine-tune themselves, rather than call a closed API, get evaluation tables showing how each checkpoint scores on MMLU, GSM8K, and HumanEval plus Chinese-language benchmarks, so fit is documented before you commit to pulling multi-gigabyte weights. The accompanying paper and open weights make it a sensible place to start when you need a permissibly licensed model and want honest evidence of what it actually does.
Articles for this project
No articles for this project yet.
To suggest a topic or contribute an article, contact us.
Related projects in this category
transformers
🤗 Transformers: the model-definition framework for state-of-the-art machine learning models in text, vision, audio, and multimodal models, for both inference and training.
CLIP
CLIP (Contrastive Language-Image Pretraining), Predict the most relevant text snippet given an image
MiniCPM-V
A Pocket-Sized MLLM for Ultra-Efficient Image and Video Understanding on Your Phone
generative-models
Generative Models by Stability AI
unilm
Large-scale Self-supervised Pre-training Across Tasks, Languages, and Modalities