skip to content
$man consulting --local --sovereign

consulting · local · sovereign

Take AI off someone else's server. Put it on yours.

We design, deploy, and operate private AI stacks: Mac mini + DGX, Qwen 3.8 / 27B / 70B and other open weights, for firms whose data, uptime, and product cannot depend on a frontier provider's good day.

// the problem

Cloud frontier models are a single point of failure for your product.

If the model that powers your workflow can be deprecated, repriced, region-blocked, or shut down by someone you've never met, your product, your clients, and your compliance posture are not yours.

E_DEPRECATED

They can kill your model.

A frontier provider sunsets a model version and your prompts, evals, fine-tunes, and product surface break overnight. You don't own the weights. You don't own the roadmap.

E_DOWNTIME

They can go down.

Region-wide outages, rate-limit storms, silent capacity throttling. If your product depends on a public LLM, a 4 a.m. incident on someone else's infra is your incident.

E_EGRESS

They can read your data.

Prompts, attachments, retrieval context: all flow through a third party's pipeline. For law, accounting, health, defense, and finance that is often a non-starter.

E_PRICING

They can change the price.

Token prices double. New "reasoning" tiers get rolled out. Your unit economics were built on a tariff that no longer exists.

E_GUARDRAILS

They can change the rules.

Silent safety policy updates, refusal drift, region blocks, content classifier changes. Your workflow worked yesterday and is "blocked" today.

E_LOCKIN

They can hold your product hostage.

If the model that powers your product can be turned off by someone else, you don't have a product. You have a lease.

We are not anti-cloud. We are pro-sovereignty. The point isn't to swear off frontier models. It's to make sure your business doesn't go dark when theirs does.

// the blueprint

A private inference stack you can describe on one slide.

Two Mac minis and one DGX-class box. Wired into your VLAN. The agents run on the Macs; the weights live on the DGX. Nothing leaves the building.
YOUR NETWORK / VLANMac Mini #1IntakeResearchAgents(LAN)Mac Mini #2ReviewerWriterWeightsBindNVIDIA DGXQwen 3.8Qwen 3-27bQwen 3-70bYour-FT32 GB500 GBM-series Processor32 GB500 GBM-series ProcessorSingle-TenantLocal GPU PoolEgress: 0 B / sUptime: 47d+Public Cloud: Optional, Never Required

// the mac minis

32 GB · 500 GB · M-series. Quiet. Cheap. Boring.

The Mac mini is the orchestrator and the agent runner. It hosts the supervisor loop, the tool-use harness, the retrieval index, the audit log, and the dashboards. It's the part of the stack you talk to and the part you replace if it breaks.

// the dgx

DGX-class box. The weights live here.

The DGX (or a single-tenant GPU pool in your cloud of choice) is where inference actually happens. Qwen 3.8 for fast cheap routing. 27B for the real work. 70B when you need the headroom. Your own fine-tunes when you need a model that already knows your business.

// the wire

Agents on the mini. Inference on the box. Zero public cloud.

The Mac minis talk to each other over the LAN. The Mac minis talk to the DGX over a private, auth'd endpoint. Nothing leaves the VLAN unless you flip a switch, and the switch is in code, not in a vendor's roadmap.

// open weights, on your metal

We don't sell you one model. We pick the right model for the job.

Open weights moved past the "good enough" line in 2025. The right answer for a routing call is not the right answer for a contract review. We carry a real menu.

3.8B

qwen 3.8

routing · classification · extraction · cheap loops

27B

qwen 27b

real work · writing · analysis · review

70B+

qwen 70b

headroom · hard reasoning · long context

custom

your weights

fine-tuned on your data · your tone · your schema

The same menu applies to Llama, Mistral, DeepSeek, Phi, and any custom-licensed weights you bring. We pick on benchmark, latency, cost, and most importantly, what your data wants.

// who this is for

If any of these are true, you don't belong on a public LLM.

The common thread: your data has a regulatory home, a contractual home, or a competitive-advantage home. That home is not 'an API call to a frontier lab.'

accounting

Accounting & CPA firms

Engagement letters, workpapers, client PII, draft audit notes. Your data has a regulatory home that is not "an API call to a frontier lab."

legal

Law firms & in-house counsel

Privileged matter, M&A diligence, contract review pipelines. The model that reads your brief cannot live on someone else's disk.

health

Health & life sciences

PHI, clinical notes, claims, prior auth. HIPAA doesn't blink at "but it's behind a login." Run the weights where the records live.

finance

Wealth, insurance & fintech

Customer portfolios, policy language, fraud signals. A private inference tier is the difference between a defensible product and a marketing claim.

public

Defense, gov & public sector

FedRAMP / IL5-style requirements, classification, citizen data. We design the on-prem envelope and the dedicated cloud envelope both.

smb

Small & medium businesses

You don't need a platform team to use AI well. You need a 32 GB Mac mini, a small box of GPUs, and the right glue.

// engagement tracks

Three ways to bring AI inside your perimeter.

Pick one. Or stitch them together. Either way, you end up with a system you own, not a relationship you depend on.

TRACK_LOCAL

On-prem. Inside your firewall. Inside your control.

Local setup

We spec, procure, rack, image, and ship a complete local inference stack: Mac mini as the orchestrator, DGX-class hardware as the weight host, agents wired to local endpoints only. Nothing leaves the building unless you turn the knob.

  • Hardware bill of materials + vendor list
  • Network + VLAN plan, including air-gapped option
  • Local LLM runtime (vLLM / llama.cpp / Ollama) + OpenAI-compatible API
  • Agent harness that targets local endpoints first, cloud only when you ask
  • Backup, monitoring, and a "what to do at 3 a.m." runbook

TRACK_DEDICATED

Single-tenant. Your weights. Your region. Your SLA.

Cloud-hosted dedicated

If you want the latency and ops story of the cloud without sharing a tenant with the public, we stand up dedicated inference: your model, your weights, your VPC, your access. No cross-tenant noise, no surprise fine-tunes, no deprecation Sunday.

  • Dedicated GPU pool on your preferred cloud
  • BYO-weights: open models, your fine-tunes, or both
  • Private networking, SSO, audit log streaming
  • Capacity commitments + per-token cost guardrails
  • Hand-back: we train your team to operate it without us

TRACK_TRAIN

Post-market training on your data. Models that speak your business.

Training & customization

Open weights are not the finish line. They're the floor. We take Qwen 3.8, 27B-class, and other strong open models, and we fine-tune them on your domain: your tone, your schemas, your policy, your prior work. The result is an SLM or LLM that behaves like a senior analyst who has read every document in your archive.

  • Data audit + curation for safe training
  • Continued pretraining, SFT, DPO, or RLAIF. Pick the lever.
  • Eval suite tied to your real tasks, not a public benchmark
  • Quantized, packaged, deployable to your hardware
  • Versioning + rollback, with a paper trail for your compliance team

// the discipline

SLMs and LLMs are not a side project for us. They're the whole practice.

We help small and medium businesses put AI at the core of how they run, not bolted on as a chat widget, but built into the workflows that already pay your bills.

$ cat expertise.txt

open weights
Qwen 3.8 · 27B · 70B · Llama · Mistral · DeepSeek · your custom
runtimes
vLLM · llama.cpp · Ollama · TGI · SGLang · TensorRT-LLM
hardware
Mac mini (M-series) · DGX · consumer GPUs · dedicated cloud GPU pools
training
LoRA · QLoRA · DPO · continued pretraining · reward modeling
agents
tool-use · retrieval · multi-step planning · browser + shell + API
ops
private networking · SSO · audit logs · cost guardrails · incident response

Most firms are trying to bolt AI onto a workflow designed for humans. That's why the demo was impressive and the rollout stalled.

We help you do the opposite. retool the business around what an AI can actually do well, then build the human loop on top of that. The result is fewer manual steps, faster turnaround, and a unit cost that improves every quarter instead of tracking the public-model price sheet.

For a 12-person accounting firm that looks like: intake agent → workpaper drafter → reviewer (human) → client-ready output, running on a single Mac mini. For a 50-attorney firm it looks like a fleet of agents on a few Mac minis wired into your DMS, with a DGX box down the hall. The shape changes; the principle doesn't.

// the question

If the model that powers your product can be turned off by someone you've never met, you don't have a product. We help you get to a model you control.

$empowered.guru --book-sovereign-session

private ai · your stack · your call

Book a sovereign-AI working session.

Bring your stack, your data shape, your regulatory constraints, and your timeline. We'll come back with a hardware spec, a model shortlist, and a week-one plan. Not a sales deck.