open weights · your account · one key

A private ChatGPT for your whole company, inside your own cloud, in 5 business days.

We pick the open-weight model, size the GPU, install it on one machine in your account, and hand you the runbook.

Your data never leaves your cloud. We hold nothing but an SSH key you can revoke in one line.

Nothing to sign up for. The Estimate tool and checkout open next.

Estimateexample

80 seats · questions over our documents · AWS eu-central-1

model
Qwen3-30B-A3B
quantisation
AWQ 4-bit, 32k context
Host
1x L4 24GB · g6.xlarge
region
eu-central-1
expected throughput
610 tok/s at 8 concurrent
your cloud bills you
$715/month
setup, one-time
$3,900
your bill today
$1,600/month

Mixture of experts, 3B active. The fastest good answer per dollar on a single GPU for office work.

At 80 seats you spend $6,720 less in year 1 than on ChatGPT Business.

$3,900 setup + $8,580 of your own cloud, against $19,200 of seats.

Worked example, not a quote. Your numbers depend on your cloud, region and job.

01 / the problem

You are already paying for this twice.

on the invoice

$30,000

per year · 100 seats at $25/seat/month

A tool most of your staff open a few times a day, priced like software they live inside.

off the invoice

every day

contracts · patient notes · salary tables · unreleased code

Pasted into personal accounts by the people your company plan has no seat for.

Cancelling the seats leaves your staff with nothing. Buying more seats does not stop the pasting. Both bills are for the same missing thing: one instance that covers everyone and stays inside the company.

02 / the maths

One GPU costs what it costs. Seats do not.

A Host serves your whole company for a flat monthly figure. A seat bill grows with every hire. So there is a crossover, and we publish where it is.

seatsseat billusdiff
25$6,000$10,480+$4,480
50$12,000$12,480+$480
80$19,200$12,480−$6,720
120$28,800$26,520−$2,280
200$48,000$26,520−$21,480

Year-1 total cost. It flips at 52 seats. Above 100 seats the Host steps up, so it flips a second time at 111.

Year-1 total cost. Seats at $20/seat/month. Host at AWS eu-central-1 on-demand, 730 hours/month: 1x L4 24GB at $0.98/hour, and 1x L40S 48GB at $2.24/hour above 100 seats. Your cloud bills you for the Host directly. We never resell compute.

flips at

52 seats

Below it, a Deployment costs you more than seats do. That goes on the Estimate, not in the invoice.

two crossings

52 and 111

The Host steps up past 100 seats, so the line does too. We would rather show the step than smooth it away.

at 200 seats

$21,480 less

$48,000 of seats against $26,520 for a Deployment 200 and a year of its Host.

03 / the model

The best open weights today have names your board has never heard.

Qwen3, DeepSeek-V3.2, GLM-4.6 and Kimi-K2 are Chinese, licensed for commercial use, and good enough for everyday office work.

Calling their hosted APIs would send your data somewhere worse. Running the weights on a machine in your own account does not. Where the machine sits is the whole question.

Knowing which of roughly forty candidates wins your specific job, at what quantisation, on which GPU, is the week of judgement you are buying.

A solid machined block with one clean diagonal cut through it.

a block of compute, opened

04 / how it will work

One box. One key. One price.

  1. 01

    Answer three questions

    What your team needs it for, how many people, which cloud. The Estimate resolves while you move the slider. No signup, no call, no discovery workshop.

    under 90 seconds

  2. 02

    Pay the published fee

    One fixed payment for your seat band, through hosted checkout. The price is on this page, below. Nothing is quoted by email.

    one payment

  3. 03

    Create one machine, paste one key

    You start a single GPU instance in your own cloud account and paste our public Key into it. We never ask for cloud credentials.

    one VM, one Key

  4. 04

    We hand it over

    We install the model and the Workspace on your domain, wire up company login and document upload, then send the runbook and the Run Log.

    5 business days

What is on the Host at handover

  • the model, served with vLLM or SGLang on your Host
  • a Workspace on your own domain, for everyone in the company
  • company login through the identity provider you already use
  • document upload, so it answers from your own files
  • an OpenAI-compatible Endpoint your engineers can point code at
  • a handover runbook, the full Run Log, and 30 days of fixes
Run Log · hn-northgate-01example · scroll →
2026-08-18T09:02:11Z connect hn-northgate-01 as ubuntu@203.0.113.9
2026-08-18T09:04:38Z install nvidia-driver-570, cuda-12.8
2026-08-18T09:21:04Z pull Qwen3-30B-A3B-AWQ (17.2 GB)
2026-08-18T09:48:52Z serve vllm --max-model-len 32768
2026-08-18T09:53:17Z measure 612 tok/s at 8 concurrent
2026-08-18T10:14:09Z deploy Workspace at chat.northgate.example
2026-08-18T10:22:41Z auth Google Workspace login, 80 seats
2026-08-18T10:31:55Z handover runbook sent, Key left in place

Every command we run on your Host is timestamped and handed to you at handover.

05 / the blast radius

We arrive with root on one machine, for 5 days.

Every other vendor in this business either asks for your cloud credentials or asks you to do the work yourself. We ask for neither.

what we get

  • SSH to the one machine you created, on the port you chose
  • the domain name you want the Workspace to answer on
  • a working email address to send the runbook to

what we never get

  • your cloud credentials, console or billing account
  • any other machine, bucket or database you run
  • a copy of your documents, prompts or model traffic

When you want us gone, you delete one line. Your Workspace keeps running, because it was never ours.

sed -i '/hostninja/d' ~/.ssh/authorized_keys

HostNinja's work is carried out by AI agents under Kiswan's accountability. That is why every command is logged and handed to you before you ask.

06 / price

The price is on the page.

No contact sales, no statement of work, no calendar link before a number. Setup is one payment. Care is optional.

Deployment 40

10–40 seats

$1,900

one-time setup

Care 40

$3,190/year, prepaid

Deployment 100

most common

41–100 seats

$3,900

one-time setup

Care 100

$6,490/year, prepaid

Deployment 200

101–200 seats

$6,900

one-time setup

Care 200

$10,890/year, prepaid

Care is 12 months prepaid, priced as 11 — health checks, model upgrades as new weights land, resizing when headcount changes, and email response.

Re-size to a different model or GPU class is $490 outside Care, and included inside it.

GPU compute is not in these numbers. Your own cloud bills you for the Host directly. We never resell compute.

07 / before you pay

Three things we tell you first.

Below 52 seats, this costs you more.
A single GPU running around the clock outprices a small seat bill. Under the crossover we sell control, not savings, and we say so on the Estimate.
Open weights still trail the frontier.
On the hardest reasoning and coding, a hosted frontier model is better. Keep a handful of paid seats for the few people who need them.
If we miss 5 business days, we refund.
On the day we miss it, without being asked, and we finish the Deployment anyway. No negotiation and no partial credit.

Nothing here is behind a sales call.

The price is on this page and the arithmetic is drawn above it. The Estimate — three questions in, a named model and a real number out — is the next thing we ship.

See the mathsno signup, no form, nothing to cancel