Interactive Archive — Literature Review
Research Archive
Scroll position drives focus strictly at the top bar. Mouse hovering highlights tiles without disrupting scroll focus.
Top Scroll Focus (Paper #1 of 10)Scroll Driven Sizing
2017Transformer Architecture & Attention MechanicsReproduced
Attention Is All You Need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Łukasz Kaiser, Illia Polosukhin • NeurIPS 20172023Small Language Models & Synthetic DataReproduced
TinyStories: How Small Can Language Models Be?
Ronen Eldan, Yuanzhi Li • arXiv Preprint (Microsoft Research)2021Parameter-Efficient Fine-Tuning (PEFT)Reproduced
LoRA: Low-Rank Adaptation of Large Language Models
Edward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, Weizhu Chen • ICLR 20222024Mixture-of-Experts (MoE) & Systems
DeepSeek-V3 Technical Report
DeepSeek-AI Team • Technical Report2015Model Compression & DistillationReproduced
Distilling the Knowledge in a Neural Network
Geoffrey Hinton, Oriol Vinyals, Jeff Dean • NIPS 2014 Deep Learning Workshop2024Small Language Models & Token Mixtures
SmolLM2: Data Mixture and Scaling for Small LMs
Hugging Face SmolLM Team • Hugging Face Technical Report2023Synthetic Data & Coding LLMs
Phi-1: Textbooks Are All You Need
Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, et al. • arXiv Preprint (Microsoft Research)2020Scaling Laws & Empirical Optimization
Scaling Laws for Neural Language Models
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, et al. • arXiv Preprint (OpenAI)2022Compute-Optimal Scaling (Chinchilla)