48 to 64 GB unified memory
Desk
from $3,000
hardware about $2,000 at cost + $1,000 build fee
A solo practice or a front office. The receptionist, the intake form, the Monday brief.
- Machine
- Mac mini, M4 Pro
- Runs
- 27B class
- Serves
- 1 to 3 people at once
Private AI kits · ships anywhere in the USA
A machine we build, harden, and load with an open-weight model and an assistant layer, then ship to your door. Plug in power and ethernet and your company has its own AI on its own network. Add the agents from the catalog and it answers the phone, too.
loading model into memory · once, at boot
Three kits
Every kit is the same software on a different amount of unified memory. More memory means a bigger model, more people at once, or both. Hardware figures are August 2026 street prices; the written quote fixes them.
48 to 64 GB unified memory
from $3,000
hardware about $2,000 at cost + $1,000 build fee
A solo practice or a front office. The receptionist, the intake form, the Monday brief.
128 GB unified memory
from $3,700
hardware about $2,200 to $3,700 at cost + $1,500 build fee
A shop, a clinic, a firm. Several agents running on the same box, client files never leaving it.
256 to 512 GB unified memory
from $9,500
hardware about $7,000 to $10,000 at cost + $2,500 build fee
Multi-location, regulated, or simply done with cloud terms. Frontier-adjacent answers, on premises.
Agents installed and ready
Any agent from the catalog, wired into your tools and running on the box before it ships, at its Quick Win price: $3,500 to $6,500 each. Desk kit plus the Intake agent is the $6,500 on the front page.
Care
Box Care, $300 a month, optional on a bare kit: model refreshes, security updates, a monthly backup check, and a person to call. Once agents are on the box the standard Care plan applies, $600 to $1,500 a month, and it is required, because agents break when vendors change things.
What arrives
It is a computer you own, running software you can inspect, with nothing on it that reports to anyone. The build is what you are paying for; here is what the build is.
Ollama or llama.cpp, MLX on Apple silicon. An OpenAI-compatible endpoint on your LAN, so anything that talks to the cloud today can talk to the box instead.
Hermes or OpenClaw, your choice. Tools, memory, and skills, running against the local model. A web chat for the team ships configured.
A current open-weight release sized to the memory: Qwen, Llama, Gemma, GPT-OSS. Chosen at install, swappable later, never pinned to a version on this page.
Full-disk encryption, firewall default-deny, no telemetry, local accounts only, automatic security updates, nightly on-box backups to a second volume.
Tailscale, installed but off. It turns on when you say so, for the session you say so, and the log shows every time it did.
One printed page in the box: what it is, how to restart it, how to revoke us, who to call. Written for whoever is there on a Saturday.
How it arrives
The clock is real: five business days on the bench, one or two in transit, and a thirty-minute call at your end. Agents, if you add them, follow the same ten-day clock as every build.
A thirty-minute call or the form below. We confirm the kit, the model, and whether agents go on it. You get a written quote with the hardware receipt price fixed.
Hardware ordered to our bench, imaged, hardened, model loaded, twenty-four hours of burn-in, and the security note written.
UPS, signature required, anywhere in the United States. Serial number and disk encryption key sent to you separately, never in the box.
Power and ethernet. A thirty-minute setup call puts it on your network and your team in the chat. It is live that afternoon.
Any agent from the catalog wired into your phone, calendar, CRM or POS, ten business days, the same as a cloud build. Same audit, one line shorter.
Terms, in plain language
Before you order
Because a markup on a Mac is not a service. You see the receipt, the manufacturer warranty is in your name, and our fee is the work: imaging, hardening, model tuning, the runbook, and the call.
For the work these agents do, reading a call, filling a form, drafting a follow-up, reconciling two lists, it is more than enough and it never sends anything out. For open-ended research and long reasoning the frontier is still better, which is why most companies run hybrid.
The box does not notice. Chat, the API, and any agent that only touches local systems keep running. Agents that talk to a cloud calendar or a phone line wait and retry.
Yes. The gateway routes anything you mark as sensitive to the box and everything else to GPT, Claude, Grok or Gemini on your keys. One audit log. The agents cannot tell the difference.
You. Local accounts you create, on your network. We hold no standing access. Remote care is off until you turn it on, and you can pull the plug on us in one click, which is literally a switch in the runbook.
Thirty minutes