Home › Blog

Just saying… some notes for the BS people out there… bragging about local usability on COT

Just saying… some notes for the BS people out there… bragging about local usability on COTs hardware (stick with your raspberry meat pies)

Three rungs, and they're far apart. Prices below are ex-VAT-ish converts at ~1.26.

**Rung 1 — where you are: £5–6k.** 4-bit on CPU+1 GPU, ~5 t/s, one stream, patient. Fine for "graph hands the LLM a finished argument to articulate." Useless for anything interactive.

**Rung 2 — model fully resident in VRAM: £30–40k.** The 2-bit quant is 239GB, so you need ~288GB of VRAM: three RTX 6000 Pro Blackwell, plus a host. That's the problem — NVIDIA's marketplace now lists the card at $13,250, up 55% in sixteen months, though street pricing has been running $8,000–9,400 if you shop. Call it £25–32k in cards, £4k host, and you jump from 5 t/s to tens. Going 4-bit doubles the card count to six and puts you at £55–65k. Also: three 600W cards plus host is ~2.2kW continuous, which is a dedicated circuit and a genuinely loud room.

**Rung 3 — run it properly: £250k+.** The native FP8 checkpoint fits on a single 8×H200 node, and with FP8 KV cache reaches the full 1M context on 8×B200. That's full precision, real concurrency, the whole context window. It's a colo decision, not a workshop decision.

**The thing that changes the maths.** MoE inference is bandwidth-bound on weights that get *reused across a batch*. A rig doing 5 t/s to one user might do 40+ t/s aggregate across sixteen concurrent requests, because the expert weights only get read once per batch. If NinjaSignal is articulating over 1.6M entities in bulk, you should be sizing for throughput, not latency — and Rung 1 hardware with proper batching gets you further than the single-stream number suggests. That's worth benchmarking before you spend Rung 2 money.

**And the thing that won't.** Token economics will never justify this. Z.ai's API undercuts your electricity bill. The only argument for owning any of it is that client adjacent and NinjaSignal data can't leave the building — which is a real argument, and it's the one to make explicitly rather than dressing it up as cost saving. If sovereignty is the driver, also price dedicated single-tenant GPU rental with a zero-retention contract: £2–5k/month, no capex, no 2.2kW in Watton, and you can stop.

Let the last part sink in…

The Probably Fine Daily

Threat intelligence every morning — new victims, new groups, what matters, in plain English. Free, with receipts.

Subscribe to the Daily →

View the original on LinkedIn ↗

← All writing