uncloseai.
How We Run Inference
How we run inference
Optional reading, for anyone reproducing our setup or contributing GPU time.
Topology
Two consumer cards, one model. Either hostname gives you Qwen3.6-27B:
| Hostname | Card | Model |
|---|---|---|
hermes.ai.unturf.com | RTX 4090 (Ada) | Lorbus/Qwen3.6-27B-int4-AutoRound |
qwen.ai.unturf.com | RTX 3090 (Ampere) | Lorbus/Qwen3.6-27B-int4-AutoRound |
Yes, hermes.ai.unturf.com serves Qwen, deliberately. Hostnames are routing labels; always read id from /v1/models.
One port per model identity. Qwen binds 18888, Hermes 18889. A hostname whose model is down gets a 502 instead of silently answering as whatever else is up.
vLLM Setup: Qwen3.6-27B on a 24 GiB card
TL;DR: INT4 AutoRound quant, fp8 KV cache, vLLM 0.26. Fits 64k context: one full-length sequence plus a few short chats batched beside it.
- Model:
Lorbus/Qwen3.6-27B-int4-AutoRound, 17.45 GiB resident. AWQ builds run 1.5 GiB heavier, which is our difference between fitting 64k and not. - fp8 KV: halves KV memory. Works on Ampere too: it is fp8 storage, dequantized to fp16 on read.
--max-model-len 65536is a hard floor; agent frameworks refuse endpoints under 64k.--max-num-seqs 8. Measured concurrency is 2 to 3; unused batch capacity costs KV pool. Lowering seqs does not enlarge our pool (util × VRAM − weights − activations).--gpu-memory-utilization 0.975on our 3090 (64k check passes by ~10 MB);0.97on our 4090, which also drives a desktop.--compilation-config {"max_cudagraph_capture_size":64}(default 224) frees hundreds of MB for KV.--tool-call-parser qwen3_xml, notqwen3_coder, which mangled this template's<function=...><parameter=...>calls and tripled empty responses.--enable-prefix-caching --enable-prompt-tokens-details: repeated system prompts hit cached KV, andusagereportscached_tokens.
Off on purpose: speculative decoding (+53% single-stream, 4x worse under 8 concurrent agents since drafted tokens reserve KV), MTP (no draft heads fit at 64k), KV offload to CPU (engine asserts under 5+ agents), CPU weight offload (~5x slower).
sudo apt-get install gcc python3-dev python3-venv ninja-build
python3 -m venv ~/vllm-venv
source ~/vllm-venv/bin/activate
pip install --upgrade pip vllm
vllm serve Lorbus/Qwen3.6-27B-int4-AutoRound \
--host 127.0.0.1 --port 18888 \
--max-model-len 65536 \
--max-num-seqs 8 \
--gpu-memory-utilization 0.975 \
--kv-cache-dtype fp8 \
--dtype auto \
--quantization auto_round \
--enable-prefix-caching \
--enable-prompt-tokens-details \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--compilation-config '{"max_cudagraph_capture_size":64}'
Our systemd unit adds:
[Service]
Environment=CUDA_HOME=/usr/local/cuda-12.8
Environment=PATH=/usr/local/cuda-12.8/bin:/usr/local/cuda-12.8/nvvm/bin:/usr/bin:/bin
Environment=CPATH=/usr/local/cuda-12.8/include
Environment=LIBRARY_PATH=/usr/local/cuda-12.8/lib64
Environment=LD_LIBRARY_PATH=/usr/local/cuda-12.8/lib64
Environment=PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
TimeoutStartSec=3600
Restart=always
MemoryMax=16G
OOMScoreAdjust=-500
Why ninja and our CUDA env: vLLM 0.26 JIT-compiles fp8 kernels through FlashInfer at warmup, which needs ninja, nvcc with its nvvm/bin siblings, headers and -lcudart. Missing any one kills our engine at warmup. Single-quote JSON args in a unit file; systemd strips bare double quotes.
vLLM Setup: Hermes-3-Llama-3.1-8B
Our previous production model, kept as a one-line swap. solidrust/Hermes-3-Llama-3.1-8B-AWQ: ~6 GiB weights, Marlin INT4 kernels native on Ampere and Ada. (FP8-Dynamic quants token-loop on a 3090; sm_86 has no fp8 tensor cores.)
--kv-cache-dtype float16must match Marlin's compute dtype;bfloat16crashes attention withquery and key must have the same dtype.--gpu-memory-utilization 0.72leaves ~4.6 GiB for an on-box F5-TTS co-tenant.--max-num-seqs 32. vLLM's default 256 OOMs during sampler warmup on 24 GiB.--tool-call-parser hermesmatches Hermes-3's<tool_call>{...}</tool_call>format.
vllm serve solidrust/Hermes-3-Llama-3.1-8B-AWQ \
--host 127.0.0.1 --port 18889 \
--max-model-len 64000 \
--max-num-seqs 32 \
--gpu-memory-utilization 0.72 \
--kv-cache-dtype float16 \
--dtype auto \
--quantization awq_marlin \
--enable-auto-tool-choice \
--tool-call-parser hermes
Proxy Setup
An edge Caddy terminates TLS and rate-limits, then forwards over TLS to a Caddy beside each GPU, which proxies to vLLM on loopback. vLLM never binds a public interface.
keepalive off on every hop that reaches vLLM. A pooled connection can be half-open; Go will not retry a POST on it, so our request vanishes and our client hangs forever. A fresh handshake is trivial next to a multi-second generation.
hermes.ai.unturf.com {
rate_limit {
zone hermes_public {
key {remote_host}
events 3
window 1s
}
}
reverse_proxy https://<gpu-host> {
header_up X-Real-IP {http.request.remote.host}
header_up X-Forwarded-For {http.request.remote.host}
transport http {
keepalive off
dial_timeout 5s
response_header_timeout 1260s
}
}
}
Access logs delete Authorization, Cookie and api-key headers via Caddy's format filter.
Model Discovery
Swagger docs at /docs: hermes.ai.unturf.com/docs, qwen.ai.unturf.com/docs.
curl https://hermes.ai.unturf.com/v1/models
curl https://hermes.ai.unturf.com/version
{
"object": "list",
"data": [
{
"id": "Lorbus/Qwen3.6-27B-int4-AutoRound",
"object": "model",
"owned_by": "vllm",
"max_model_len": 65536
}
]
}
Pass id as your model name. We omit --served-model-name: an alias hides which checkpoint you have.
Thinking mode: roughly 8x slower. Our clients send "chat_template_kwargs": {"enable_thinking": false} by default.
Rate Limiting & Access
hermes.ai.unturf.com: public, 3 requests per second per IP.qwen.ai.unturf.com: API key required; anonymous traffic gets 403.
Next Steps
Ready to add text-to-speech to your application?