A curated Python-based GitHub list of efficient LLM techniques with a broad topic taxonomy and links to sub-pages. The repository has 2031 stars and 169 forks; last push 2025-06-17.
Collecting history — the radar snapshots this repo daily. The trend line appears after 3 days of data (1 so far).
What it is
A curated list for Efficient Large Language Models.
How it works
The README shows a structured set of sub-pages under topics such as Network Pruning / Sparsity, Knowledge Distillation, Quantization, Inference Acceleration, Efficient MOE, Efficient Architecture of LLM, KV Cache Compression, Text Compression, Low-Rank Decomposition, Hardware / System / Serving, Efficient Fine-tuning, Efficient Training, and Survey.
Getting started
"Please check out all the papers by selecting the sub-area you're interested in. On this main page, only papers released in the past 90 days are shown." The repository provides a generator instruction for contributions:
- "If you'd like to include your paper, or need to update any details such as conference information or code URLs, please feel free to submit a pull request. You can generate the required markdown format for each paper by filling in the information in
generate_item.pyand executepython generate_item.py."
Recent releases
RELEASES (latest 0):
- none
Traction
stars_7d: 0 stars_1d: 0
Behind the repo
Linked to a broader list structure and contains a section for updates, including a May 29, 2024 update and a April 15, 2025 update linking to another list for Efficient Reasoning Models.
Caveats
license: none listed created: 2023-05-22 last_push: 2025-06-17 language: Python open_issues: 11 forks: 169 stars: 2031






