uncloseai.

How We Run Inference

How we run inference

Optional reading, for anyone reproducing our setup or contributing GPU time.

Topology

Two consumer cards, one model. Either hostname gives you Qwen3.6-27B:

HostnameCardModel
hermes.ai.unturf.comRTX 4090 (Ada)Lorbus/Qwen3.6-27B-int4-AutoRound
qwen.ai.unturf.comRTX 3090 (Ampere)Lorbus/Qwen3.6-27B-int4-AutoRound

Yes, hermes.ai.unturf.com serves Qwen, deliberately. Hostnames are routing labels; always read id from /v1/models.

One port per model identity. Qwen binds 18888, Hermes 18889. A hostname whose model is down gets a 502 instead of silently answering as whatever else is up.

vLLM Setup: Qwen3.6-27B on a 24 GiB card

TL;DR: INT4 AutoRound quant, fp8 KV cache, vLLM 0.26. Fits 64k context: one full-length sequence plus a few short chats batched beside it.

Off on purpose: speculative decoding (+53% single-stream, 4x worse under 8 concurrent agents since drafted tokens reserve KV), MTP (no draft heads fit at 64k), KV offload to CPU (engine asserts under 5+ agents), CPU weight offload (~5x slower).

sudo apt-get install gcc python3-dev python3-venv ninja-build
python3 -m venv ~/vllm-venv
source ~/vllm-venv/bin/activate
pip install --upgrade pip vllm

vllm serve Lorbus/Qwen3.6-27B-int4-AutoRound \
    --host 127.0.0.1 --port 18888 \
    --max-model-len 65536 \
    --max-num-seqs 8 \
    --gpu-memory-utilization 0.975 \
    --kv-cache-dtype fp8 \
    --dtype auto \
    --quantization auto_round \
    --enable-prefix-caching \
    --enable-prompt-tokens-details \
    --enable-auto-tool-choice \
    --tool-call-parser qwen3_xml \
    --compilation-config '{"max_cudagraph_capture_size":64}'

Our systemd unit adds:

[Service]
Environment=CUDA_HOME=/usr/local/cuda-12.8
Environment=PATH=/usr/local/cuda-12.8/bin:/usr/local/cuda-12.8/nvvm/bin:/usr/bin:/bin
Environment=CPATH=/usr/local/cuda-12.8/include
Environment=LIBRARY_PATH=/usr/local/cuda-12.8/lib64
Environment=LD_LIBRARY_PATH=/usr/local/cuda-12.8/lib64
Environment=PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
TimeoutStartSec=3600
Restart=always
MemoryMax=16G
OOMScoreAdjust=-500

Why ninja and our CUDA env: vLLM 0.26 JIT-compiles fp8 kernels through FlashInfer at warmup, which needs ninja, nvcc with its nvvm/bin siblings, headers and -lcudart. Missing any one kills our engine at warmup. Single-quote JSON args in a unit file; systemd strips bare double quotes.

vLLM Setup: Hermes-3-Llama-3.1-8B

Our previous production model, kept as a one-line swap. solidrust/Hermes-3-Llama-3.1-8B-AWQ: ~6 GiB weights, Marlin INT4 kernels native on Ampere and Ada. (FP8-Dynamic quants token-loop on a 3090; sm_86 has no fp8 tensor cores.)

vllm serve solidrust/Hermes-3-Llama-3.1-8B-AWQ \
    --host 127.0.0.1 --port 18889 \
    --max-model-len 64000 \
    --max-num-seqs 32 \
    --gpu-memory-utilization 0.72 \
    --kv-cache-dtype float16 \
    --dtype auto \
    --quantization awq_marlin \
    --enable-auto-tool-choice \
    --tool-call-parser hermes

Proxy Setup

An edge Caddy terminates TLS and rate-limits, then forwards over TLS to a Caddy beside each GPU, which proxies to vLLM on loopback. vLLM never binds a public interface.

keepalive off on every hop that reaches vLLM. A pooled connection can be half-open; Go will not retry a POST on it, so our request vanishes and our client hangs forever. A fresh handshake is trivial next to a multi-second generation.

hermes.ai.unturf.com {
    rate_limit {
        zone hermes_public {
            key {remote_host}
            events 3
            window 1s
        }
    }
    reverse_proxy https://<gpu-host> {
        header_up X-Real-IP {http.request.remote.host}
        header_up X-Forwarded-For {http.request.remote.host}
        transport http {
            keepalive off
            dial_timeout 5s
            response_header_timeout 1260s
        }
    }
}
        

Access logs delete Authorization, Cookie and api-key headers via Caddy's format filter.

Model Discovery

Swagger docs at /docs: hermes.ai.unturf.com/docs, qwen.ai.unturf.com/docs.

curl https://hermes.ai.unturf.com/v1/models
curl https://hermes.ai.unturf.com/version
{
  "object": "list",
  "data": [
    {
      "id": "Lorbus/Qwen3.6-27B-int4-AutoRound",
      "object": "model",
      "owned_by": "vllm",
      "max_model_len": 65536
    }
  ]
}

Pass id as your model name. We omit --served-model-name: an alias hides which checkpoint you have.

Thinking mode: roughly 8x slower. Our clients send "chat_template_kwargs": {"enable_thinking": false} by default.

Rate Limiting & Access

Next Steps

Ready to add text-to-speech to your application?

📖 Read the Text-to-Speech documentation →