Home | Markdown | Gemini | Microblog | Street photography | Wishlist

Running LLMs by myself - Part 1: Hyperstack and vLLM



Published at 2026-10-06T23:20:00+03:00

The gt calculator was built almost entirely with self-hosted LLMs running on rented Hyperstack VMs, and I ended that post with a promise: "I will write another blog post at some point about my setup and what I learned from self-hosting models on Hyperstack." This is that post.

2026-06-01 gt calculator - a calculator built with local LLMs

Here's my setup, what I learned about inference, and the numbers from using it. I still haven't bought the hardware. This is the first of two parts. The second one is about the pi coding agent, the harness I use with these models.

2026-10-06 Running LLMs by myself - Part 1: Hyperstack and vLLM (You are currently reading this)
2026-10-07 Running LLMs by myself - Part 2: The pi coding agent

The hypr logo

Table of Contents




Why rent instead of buy



The main motivation was a question I wanted to answer before spending money: should I buy hardware to run my own LLMs? A new RTX 5090 with 32 GB VRAM costs a few thousand dollars, and a 128 GB DGX Spark is now reported at $6,950, up from $4,699 in February and $3,999 at launch (thanks to the RAM-ocalypse, NVIDIA now even sells a 64 GB DGX Spark, starting at $4,999). Anyways, before putting either on my desk, I wanted to know what it feels like to have a 27B model (and the occasional 120B one) as a daily coding and planning partner. So I rented instead: Hyperstack VMs with A100 80 GB GPUs, spun up when I need them, torn down when I don't.

Hyperstack

A few months in, the setup has grown into a small toolchain I call hypr. One Ruby script manages the whole lifecycle (create the VM, open a WireGuard tunnel, start vLLM with the model of choice), and I can run two VMs at the same time, each with its own model and its own pi coding agent in a tmux pane. For example, Qwen3.8 27B works on one project in pane 0 while Gemma 4 31B works on another in pane 1, and can get a 2nd opinion from a different model this way. But often, I simply only run Qwen3.8 on only one VM without the second since that's enough for most of my uses since that model is now soo good that it can do both: Coding and has a lot of good general knowledge. So there is no need anymore to have two models running side-by-side.

hypr on GitHub
Pi coding agent

The setup at a glance



Everything starts on my laptop. ruby hyperstack.rb --vm 1 create (or --vm both for two VMs in parallel) provisions a Hyperstack VM (Ubuntu 24.04 with CUDA and Docker preinstalled, one A100 80 GB PCIe GPU), sets up a WireGuard tunnel to it, and starts the model in a vLLM Docker container. About ten minutes later the model is serving on an OpenAI-compatible API, reachable over the tunnel.

The setup: pi on the laptop, vLLM on two Hyperstack VMs, WireGuard in between

The WireGuard tunnel is the interesting bit. A single wg1 interface on my laptop (Fedora Linux, by the way!) routes traffic to both VMs. The firewall on the Hyperstack VMs only lets in SSH and the WireGuard port, so the vLLM port (11434) is only reachable through the tunnel. There's no other authentication: no proxy, no load balancer, no API keys. The API isn't exposed to the internet at all, and the tunnel is the security boundary.

hypr is a Ruby script with a small TOML config per VM. The commands that matter:


What decides if this works



Before the walkthrough, a short glossary of the things that matter when you run LLMs yourself. The bigger topics get their own section below, come back later in this post, or are covered in part 2:


The next section explains how inference works: VRAM, the KV cache, prefill, decode and the runtime. MoE, quantization and reasoning come with the models after that, and the harness topics (system prompt, tool calling) get their own post, the second part of this series.

How the inference works



I wanted to understand where the time and memory go. This is what I learned.

CUDA: how vLLM uses the GPU



CUDA is NVIDIA's software platform for running calculations on its GPUs. In my setup, the CPU runs vLLM's request handling and decides which work to send to the GPU. Through CUDA, it launches small GPU programs called kernels. The GPU does the model's heavy maths, using the weights and cached context in its own memory. The Hyperstack VMs I'm renting use CUDA.

CPU organises the work, CUDA launches it, GPU does the maths

I don't need to write CUDA code for this. vLLM and its libraries take care of it. The NVIDIA Container Toolkit lets the Docker container use the host's GPU and driver with --gpus all.

What vLLM is



vLLM is an open-source inference engine: the program that takes a model's weights and turns prompts into tokens at high throughput. Ollama does the same job packaged for convenience: one command pulls a quantized model and it's serving. vLLM is the lower-level engine you run yourself, in Docker, with flags, and tune per model. Ollama optimizes for "make it run"; vLLM optimizes for throughput, context length, and batching, and exposes the knobs (--max-model-len, --gpu-memory-utilization, --enable-prefix-caching) so you can spend the VRAM the way your workload needs it.

An LLM request has two phases:


One request: prefill, then decode

The two phases show up as very different numbers. Prefill can process thousands of tokens per second; ordinary decode in my setup produces tens per conversation. During prefill, many tokens share the work of reading weights from GPU memory. During decode, that cost returns at each step. Batching conversations and speculative decoding (below) can get more tokens out of each pass. Prefill and any queueing affect the wait for the first token; decode affects how quickly the answer arrives after that.

One more knob that is good to know about is the temperature: it controls how adventurous the model is when it picks the next token during decode. I currently have no temperature setting in my hypr setup, though.

The key data structure is the KV cache. Full-attention layers reuse key and value tensors from earlier tokens instead of recomputing the whole history for each new token. Their cache grows with the context. My models also have layers that work differently: Qwen3.8-27B has 48 Gated DeltaNet layers with recurrent state and 16 full-attention layers; Gemma 4 mixes sliding-window and global attention. So the exact memory cost depends on the LLM architecture.

The context window is a length limit, while the cache is the memory used to serve requests within that limit. A request needs room for both its prompt and its generated answer, including thinking tokens. Spare cache memory does not raise the configured or supported context limit.

Qwen3.8-27B architecture and model card
vLLM: cache management for hybrid models

VRAM = model weights + KV/state cache pool + runtime overhead + headroom

--gpu-memory-utilization 0.92   → vLLM may use 92% of the 80 GB
--max-model-len 262144          → max context: 262,144 tokens (the "262K")

At startup, vLLM loads the weights, then budgets runtime memory and a pool of cache blocks within the allowed VRAM. That is why nvidia-smi (a CLI for NVIDIA cards) shows ~72–75 GiB "used" even when nothing is running. The pool is preallocated, not busy.

How three presets split the A100's 80 GB between weights and KV cache

Prefix caching: the big one for agentic coding



--enable-prefix-caching lets vLLM retain reusable cache blocks between requests until it needs to evict them. If a new request shares a prefix with a previous one (same system prompt, same conversation history), matching cached blocks can skip prefill and therefore speed up the total time until the first token. Evicted blocks and an incomplete final block may still need work.

For a coding agent this is important. Every turn of a coding agent session resends the system prompt plus the entire conversation. Without prefix caching, that whole context is re-prefilled on every single turn. With it, most of the prompt arrives already computed. My current sessions sit at an 80–97% hit rate (numbers below).

Prefix caching over four agent turns

vLLM's cache can retain prefixes from several conversations. Ollama has prefix reuse too.

Ollama's prefix-matching cache implementation
Ollama: Flash Attention and parallel requests
Ollama: configuring context length

Ollama is still nicer if I just want a model serving quickly. For several coding agents, I prefer vLLM's batching and cache controls.

Speculative decoding: 2.7× faster decode



There's one more decode trick, and it turned out to be the biggest speedup in this whole setup: speculative decoding. The idea is to guess the next few tokens cheaply and let the big model check all the guesses in one go.

Speculative decoding: a cheap guess of 4 tokens, verified by the big model in one pass

Here's a simplified example using greedy decoding, where the model always picks its highest-scoring token:


vLLM: speculative decoding and its guarantees

Why can this be faster? Several draft positions share a read of the main model's weights. At small batch sizes, checking those extra positions can be relatively cheap. Drafting and verification still cost time and memory, though, and rejected guesses waste work. The speedup depends on how often the guesses are accepted. Predictable bits of code, such as boilerplate and repeated identifiers, are a useful case to try.

With several agents already batched, there may be less spare capacity for speculation. So I measured both cases.

Qwen3.8's config has mtp_num_hidden_layers: 1, so the model ships its own MTP head, and vLLM supports it. In vLLM, it's one flag: --speculative-config '{"method": "mtp", "num_speculative_tokens": 3}'. I added it to the Qwen3.8 preset in hypr. It's on by default now, and --no-speculative on create or model switch turns it off again.

To see if it pays off, I benchmarked it on the A100: six Go coding prompts with up to 700 output tokens each, thinking off, first one conversation at a time and then three in parallel. I ran it without speculative decoding and with 2, 3 and 4 drafted tokens. The chart shows sampling at temperature 0.7 and top_p=0.8:

Measured decode speed and aggregate throughput with and without MTP

That's a lot more than I expected. Median decode speed for a single conversation went from 48 to 132 tokens per second with 3 drafted tokens, 2.7 times as fast. This excludes the wait for the first token. Three conversations together went from 134 to 286 tokens per second, counting all wall time including that wait. About 86% of the drafted tokens were accepted: roughly 2.6 accepted guesses plus one target-supplied token per verification step. The MTP head did well on these prompts.

I went with 3 drafted tokens for now. It had the best total throughput with three sampled conversations in this run. Four won with greedy decoding, and even with sampling its per-request decode speed was slightly higher. But its median wait for the first token rose to 1.21 seconds with three sampled conversations; with three drafts it stayed around 0.4 seconds. I only ran six prompts per case, without repeated trials, so this isn't a settled optimum.

The memory cost was about 10% of the KV cache pool: 657K down to 588K tokens, still enough for more than two full 262K contexts.

These were short prompts with thinking off. I haven't measured the speedup in long coding sessions or with thinking on. Even the baseline was faster than the ~40 tokens per second from my recorded work morning (see "The numbers" below).

Qwen3.8 FP8 configuration
vLLM's Qwen3.8 recipe, including MTP

The VMs and the models



Both VMs are n3-A100x1 flavors: one A100 80 GB PCIe, 28 vCPUs, 120 GB RAM, in a Canadian region. When the A100 flavor is sold out, I manually change the config to n3-H100x1. It also has 80 GB of GPU memory, but its hardware, price and performance differ; it is not the machine behind the A100 measurements below.

The default model on VM1 is Qwen/Qwen3.8-27B-FP8, a dense 27B model with a native 262K context window, FP8-quantized. As of writing, that is also my main model in this setup, and the throughput and startup observations below are from that setup. VM2 runs Gemma 4 31B (AWQ 4-bit) by default, so I can work on two projects in parallel with two different models and compare how they behave.

Each VM's TOML config defines named presets, so switching models does not mean reprovisioning. The memory figures below are approximate weight footprints, not total vLLM allocations. The context lengths are my configured limits, not necessarily the models' native limits:


ruby hyperstack.rb --vm 1 model switch nemotron-super stops the running container, starts a new one with the preset's flags, and waits for readiness. The weights are cached on the VM's ephemeral NVMe disk after the first download, but a switch still takes minutes (see below).

Two details:


Why starting a model takes minutes



Loading a model onto the GPU sounds like "copy 29 GiB from disk to VRAM". That part is actually fast. My provisioning timings and vLLM logs give this approximate breakdown:

Where the startup time goes: fresh VM vs container restart

On a fresh VM, hypr create took about 10 minutes in total:


The last three steps prepare the GPU work. Compilation combines and optimises operations, profiling estimates the memory needed during inference, and CUDA graph capture records launch sequences that can be replayed with less CPU overhead. My log reported a 40.9 GiB cache pool and 86 captured graphs.

A restart of the same container is faster: about 2 minutes. The weights are already on disk, and the compiled kernels come from vLLM's on-disk compile cache (0.55 seconds instead of about 40). But the CUDA graphs are captured again on every start (~1 minute), and the Python and API startup doesn't get faster either.

That's why switching models ad hoc is painful. model switch takes at least 2 minutes for a model that was already used on this VM, and more for a new one (download plus a cold compile). On top of that, every running agent session loses its prefix cache and has to prefill its whole context again. So instead of switching back and forth, I run two VMs with two different models loaded. Switching between them is then just picking the other provider in pi (Ctrl+L), and it's instant.

MoE vs dense, and quantization



A dense model uses all of its parameters for every token. A mixture-of-experts (MoE) model is split into many "experts", and a small router picks only a few of them per token. So only a small active set does the work, e.g. 3B active out of 35B total.

Dense vs mixture-of-experts: active parameters versus total weight storage

Reading fewer weights per token can make small-batch decode faster than on a dense model of the same total size. How much faster depends on the routing, kernels and batching, too. In my fully GPU-resident setup, all experts stay in VRAM because the router can choose different ones for the next token. Offloading experts to CPU memory is possible in other setups, but changes the performance trade-off.

Several presets are MoEs (Nemotron-3-Super, qwen36-35b-a3b, qwen3-coder-30b). Only the active parameters do the work per token, which is why Nemotron can be 120B total with 12B active and still decode at a usable speed. But in this setup all 120B parameters still have to fit into VRAM, and that's why it barely fits once the context is capped. The qwen36-35b-a3b preset only fits into ~23 GiB because it's also 4-bit quantized.

Quantization is how the rest of the list fits. FP8 stores the quantized weights in 8 bits (a 27B model is ~27 GB of weights, plus some overhead). AWQ uses 4-bit weights for the quantized layers, plus scales and any layers kept at higher precision. It reduces the footprint; the quality and speed trade-off depends on the model and kernels. FP8 on the 27B is my daily driver: good quality, 28.9 GiB of weights, and 40.9 GiB left for the KV cache. vLLM reports that as 657,281 tokens, enough for about 2.5 full 262K contexts at the same time. I haven't compared these presets against quantization-aware training (QAT) variants on my own tasks. QAT makes lower precision part of training rather than only applying it afterwards.

Why this is the daily driver



The official model card benchmarks Qwen3.8-27B against its predecessor, Qwen's own closed-weight Qwen3.7-Plus, a 30B-class competitor, and the frontier Opus 4.6 Max. Here are selected coding and general-reasoning rows as a chart:

Qwen3.8-27B vs Qwen3.6-27B, Qwen3.7-Plus, Muse Glimmer-30B and Opus 4.6 Max (vendor-reported)

It beats its direct predecessor on every row, and the rows that match how I use it are the interesting ones: on SWE-bench Pro (agentic coding) the 27B dense model scores 61.7, ahead of Opus 4.6 Max's 53.4 and Muse Glimmer's 51.2; on terminal coding it runs 73.0 to Qwen3.6's 63.4; on competitive coding (90.3) it even edges Opus (88.8). Those results made it worth trying as a coding partner on one rented A100.

These are vendor-reported scores, and a couple of the benchmarks are Qwen's own. For SWE-bench Pro, Qwen corrected some tasks and re-evaluated the other models with Claude Code, but retained the previously reported Opus score. That makes 61.7 versus 53.4 an uneven comparison. The table also doesn't establish the score of my FP8 checkpoint running through pi.

Qwen's benchmark table and evaluation notes

Simon Willison tried the model on a DGX Spark and an M5 Max MacBook the week it shipped. His experience of overthinking matches mine: the model defaults to xhigh reasoning effort.

Qwen 3.8 27B is excellent, but it defaults to wildly overthinking things (Simon Willison)

Reasoning: what it is and the effort levels



Reasoning models "think" before they answer. They write out their thoughts first, usually between <think> and </think> tags, and only then produce the answer or tool call. They were trained (mostly with reinforcement learning) to work through a problem step by step in that scratchpad, and for tricky tasks the answers get noticeably better.

For the inference engine, thinking tokens are completely normal output tokens. They're decoded one by one at decode speed, and they fill the KV cache like everything else. vLLM's reasoning parser (--reasoning-parser qwen3 in my presets) only splits them from the answer, so pi can show them separately. They are the greyed-out text in the pi screenshots in part 2.

Reasoning effort: more effort means more thinking tokens before the answer

Many models offer reasoning effort levels: off, low, medium, high, and sometimes xhigh. The model was trained to think shorter or longer depending on the level. How the level reaches the model depends on the model and the API:


And not every model supports it:


Here's a catch I only found while writing this post. pi shows "medium" as the thinking level in its footer. But with the qwen-chat-template compatibility setting I use, pi sends enable_thinking: true or false to vLLM without a reasoning effort. So low, medium and high all just mean "on", and Qwen3.8 then thinks at its own default, xhigh. That fits the overthinking I've been seeing. pi also asks the chat template to keep the thinking of earlier turns in the history (preserve_thinking), so the thinking keeps taking up context later, too. And for Gemma and Nemotron, my pi config marks the models as non-reasoning, so pi's level does nothing there at all. For Nemotron, that's a config gap on my side: its chat template can enable thinking independently of pi's setting. The reasoning parser only separates that output; it does not turn thinking on.

I haven't done a clean on-vs-off comparison on my own tasks yet. But the math is simple: at ~40 tokens per second, 2,000 thinking tokens are 50 seconds before the first word of the answer. For agent turns where I already know what the change should look like, that's mostly wasted time and context. Worth a dedicated experiment later. For now, I live with the default and interrupt the model when it spirals.

Pi's separate qwen compatibility path can also send reasoning_effort when configured for it. The on/off behaviour above comes from my compatibility setting.

Pi's thinking-parameter handling

Inside the VM



Once the tunnel is up, the VM is just another machine on your network, and I mostly work on it over its WireGuard address rather than the public IP:

$ ssh ubuntu@192.168.3.1
ubuntu@hyperstack1:~$ hostname
hyperstack1

The entire "AI service" from the inside is one Docker container:

$ docker ps
CONTAINER ID   IMAGE                      COMMAND                  CREATED       STATUS       PORTS     NAMES
aa1845729f93   vllm/vllm-openai:nightly   "vllm serve --model …"   2 hours ago   Up 2 hours             vllm_qwen38_27b

The interesting part of the container's log is the Engine 0 line vLLM emits every ten seconds, the same line the watch dashboard parses. This snapshot is from later in the day than the numbers further down:

$ docker logs --tail 4 vllm_qwen38_27b 2>&1 | grep "Engine 0"
(APIServer pid=1) INFO 09-13 08:56:54 [loggers.py:311] Engine 000: Avg prompt
throughput: 0.0 tokens/s, Avg generation throughput: 53.6 tokens/s, Running:
2 reqs, Waiting: 0 reqs, GPU KV cache usage: 59.3%, Prefix cache hit rate: 96.9%
(APIServer pid=1) INFO 09-13 08:57:14 [loggers.py:311] Engine 000: Avg prompt
throughput: 0.0 tokens/s, Avg generation throughput: 48.3 tokens/s, Running:
1 reqs, Waiting: 0 reqs, GPU KV cache usage: 33.3%, Prefix cache hit rate: 96.9%

Follow it live with docker logs -f vllm_qwen38_27b 2>&1 | grep "Engine 0", or watch the hardware side with nvidia-smi --query-gpu=temperature.gpu,utilization.gpu,power.draw --format=csv -l 5.

The numbers



I kept 474 log samples from about 80 minutes of work on the A100, running Qwen3.8-27B-FP8 with a 262K context limit. Here's one:

Engine 000: Avg prompt throughput: 253.8 tokens/s, Avg generation
throughput: 35.9 tokens/s, Running: 1 reqs, Waiting: 0 reqs,
GPU KV cache usage: 5.0%, Prefix cache hit rate: 81.1%

ruby hyperstack.rb watch combines these ten-second log intervals with nvidia-smi in a dashboard that refreshes every two seconds:

The hypr watch dashboard during a decode burst


The surprise was running three agents in three tmux panes against one GPU. I didn't notice a slowdown. Batching shares weight reads across conversations, although each still adds compute and cache work.

A few dashboard details caught me out. VRAM stays near 90% even when idle because vLLM reserves the cache pool. The separate KV-cache bar shows how much of that pool is in use. Neither bar tells me how close one conversation is to its 262K limit. running includes prefill as well as decode, and 100% GPU utilisation means a kernel was running throughout the sample, not that every compute unit was fully used.

NVIDIA: GPU utilisation

Short turns sometimes finished in 10–15 seconds; long answers and thinking took longer. My old Ollama setup felt slower, but its settings differed, including a 32K context limit. I wouldn't call that a benchmark.

What it costs, and do I buy the hardware?



The A100 costs $1.35/hour on demand, billed per minute. The H100 fallback is $2.50/hour. Reserved prices start at $0.95 and $1.75 respectively (September 2026). The compute alone would cost $972 for a 30-day month. Hyperstack also lists a public-IP charge of $0.00672043/hour; with one billed IP, that comes to about $976.84. I delete mine when I'm done.

The per-token calculations below use the $1.35/hour compute rate. They exclude the public IP and any additional storage.

Hyperstack pricing

That recorded morning, before speculative decoding, used about $1.78 of compute for 79 minutes and produced roughly 250K output tokens. Charging the rounded compute cost to output gives about $7.12 per million tokens, including prompt processing and time spent waiting on tools.

Speculative decoding made the tokens cheaper in my short-prompt benchmark. Counting the compute time for the whole run:


Those runs had thinking off and included the wait for the first token, but no tool calls or idle gaps between coding turns. I haven't measured the savings over a full work session yet. The VM still costs $1.35/hour either way.

Output costs: GPU rental with and without MTP, the earlier work session, and OpenRouter examples

For an API comparison, OpenRouter lists these prices for Qwen3.8 27B on 30 September 2026, per million tokens:


The chart shows their output charges; input costs come on top. My GPU bars charge all the compute time to output, including prefill and waits, but exclude the IP and storage charges. I still can't price the full work session through an API: the earlier 11M input-token estimate needs checking because current vLLM logs exclude cache hits from prompt throughput.

OpenRouter's Qwen3.8 provider prices, checked 30 September 2026
vLLM's token accounting

Buying still doesn't appeal to me:


RTX 5090 specifications
A100 specifications
DGX Spark price change, February 2026
DGX Spark 64 GB at $4,999, 128 GB at $6,950, checked 6 October 2026
DGX Spark specifications

Spark does run CUDA and vLLM, with a compatible ARM container. I could use the same Hugging Face serving approach there; it isn't limited to GGUF or LM Studio.

Running vLLM on DGX Spark

A $5,000 machine, like the 64 GB Spark, equals about 3,704 hours at the A100 compute rate. At four hours every day, that's two and a half years. The 128 GB Spark at $6,950 is about 5,148 hours, or three and a half years. All of that is before electricity, resale value and performance differences. Buying could pay off. I'm just not ready to bet on which machine I'll still want by then.

Wrapping up



I'll keep renting for now. I still want a GPU on my desk, partly for privacy and partly because I like owning the tools I use. A rented VM is still somebody else's hardware. Perhaps replacing my ThinkPad will be the excuse; perhaps I'll wait for a faster unified-memory box.

I'd keep my cloud subscriptions either way. In my daily work, Opus 5.5 handles the big, messy tasks better. The self-hosted models are useful when I'm closely involved: studying code, writing a test, making a small change, being more in-the-loop. That's how I built gt, and it's how I enjoy using them.

The provisioner, model presets and pi extensions are all here:

hypr on GitHub

Read the next post of this series:

Running LLMs by myself - Part 2: The pi coding agent

Other related posts:

2026-10-07 Running LLMs by myself - Part 2: The pi coding agent
2026-10-06 Running LLMs by myself - Part 1: Hyperstack and vLLM (You are currently reading this)
2026-06-01 gt calculator - a calculator built with local LLMs
2025-08-05 Local LLM for Coding with Ollama on macOS

E-Mail your comments to paul@nospam.buetow.org :-)

Back to the main site