Inference Stack
The reference for how AI infrastructure actually fits together
Architectures
Catalog
Where it runs
About
Back to the interactive map
Inference Frameworks
SGLang
Structured generation with RadixAttention KV-cache reuse, speculative decoding, and fast scheduling
Official docs
Used by
Text Generation
RAG
Leads to
SGLang Runtime
Ray Serve