

LLM Inference & Serving — Synthszr Charts
This category covers frameworks and engines used to run large language models in production – from inference servers to batching and scheduling layers to optimizations for latency and throughput. It includes projects like vLLM, SGLang and TGI, used to serve models behind APIs or in self-hosted data centers. They're used by teams that need to run models at scale and cost-efficiently, not just train them.
The Synthszr Charts measure momentum for this category from mentions across thousands of news and newsletter sources, recency-weighted and version-granular, so that new releases such as EAGLE 3.1 are tracked individually. The ranking is updated daily, showing which serving solutions are gaining or losing attention.
- 1vLLMvllm · 59× · Aug 13, 2026100
- 2Lightningnvidia · 4× · Aug 13, 202672
- 3Computercloudflare · 3× · Aug 07, 202650
- 4SubQsubquadratic · 18× · Aug 17, 202645
- 5TabFMgoogle · 9× · Jul 27, 202641
- 6SGLanglmsys · 11× · Aug 05, 202638
- 7Workers AIcloudflare · 2× · Aug 08, 202629
- 8Mooncakemoonshot · 2× · Jul 19, 202619
- 9ModelOptnvidia · 2× · Jul 15, 202616
- 10Yansubentoml · 3× · Jun 01, 20263
- 11TGIhugging-face · 2× · Jun 19, 20263
- 12CoreWeave Sandboxescoreweave · 3× · May 24, 20262
- 13Text Generation Inferencehugging-face · 2× · May 28, 20261
- 14Groq Llamagroq · 2× · May 17, 20261
- 15vLLM 0.20.0vllm · 4× · Apr 29, 20261
- 16TokenSpeedlightseeking · 2× · May 12, 20261
- 17Docker Model Runnerdocker · 2× · May 08, 20261
- 18HybridClawhybridai · 2× · Apr 29, 20260
- 19NEAR AI Cloudnear-ai · 7× · Apr 03, 20260
- 20IronClawnear-ai · 7× · Apr 03, 20260
- 21TensorRT-LLMnvidia · 2× · Apr 29, 20260
- 22ClawProtencent · 2× · May 04, 20260
- 23HDNA Workbenchunknown · 2× · Apr 15, 20260
- 24Crusoe Managed Inferencecrusoe · 2× · Apr 07, 20260