vLLM
5 pieces on vLLM, including 2 step-by-step guides.
Guides
Run LLMs Locally: 6 Practical Methods (Ollama, LM Studio, vLLM, Jan, llama.cpp, and llamafile)
Master 6 proven frameworks to host and run LLMs locally across Windows, macOS, and Linux with GPU acceleration, from one-click GUIs to high-throughput production servers.
Self-Host DeepSeek V4 with vLLM: 2026 Deployment Guide
Deploy DeepSeek V4 Flash or Pro with vLLM, size the GPU cluster, configure tensor parallelism, and compare self-hosting with current API prices.
News & Analysis
2.4T-Parameter Qwen3.8 Runs on One Node With vLLM and NVFP4
AWS documents a single-node serving path for Qwen3.8-2.4T-A95B on a p6-b300 instance using NVFP4 quantization, cutting TTFT by nearly 60%.
DFlash Delivers 3.92x CPU Token Throughput on Qwen3.5-9B
DFlash in vLLM v0.25.0 hits 3.92x average token throughput on a Xeon 6975P-C CPU — a 74.40% cost cut with one config flag.
PagedAttention vs RadixAttention: How LLMs Tame the KV Cache
Two architectures attack LLM KV cache inefficiency from opposite angles — memory fragmentation and redundant prefill. Here is how each works.