LLM Serving Glossary
Plain-language explanations of terminology from the world of self-hosted LLM inference and serving. Written for someone without an LLM/GPU background — each entry says what the thing is and why it matters in practice.
Organized by topic (see the sidebar), with full-text search at the top. Synonyms and abbreviations are listed in parentheses. Terms link to a primary source (paper, spec, or docs) where one exists.
Missing a term, or found a definition unclear or wrong? Every page has an “Edit page” link at the bottom that takes you straight to the source on GitHub — or open an issue. Corrections and suggestions of any size are welcome.
Everything here is CC0 / public domain: use it however you like, no attribution needed.