inference
inference related cheatsheet collection - 1 quick reference guides. Find fast, copy directly. The essential online command reference for developers and SRE.
vLLM Cheatsheet - High-Performance LLM Inference & Serving
One-page vLLM (high-performance LLM inference & serving engine) reference covering install, OpenAI-compatible server launch, host/port/API-key config, and quantization. Includes pip install vllm, vllm serve, --host/--port/--api-key. Copy-ready.