Services / 05 · Private AI
A model that runs where your documents already live.
A self-hosted language model answering questions over your own files — contracts, case files, patient records, customs paperwork — on hardware you control, with nothing sent to somebody else’s service to be answered. It will sometimes be wrong, and it does not make you compliant. Both of those are on this page rather than in the small print.
What is delivered
Four parts, and the one in the middle is where the work actually is.
Standing a model up is the easy half and any competent engineer can do it in an afternoon. Making it answer usefully from your documents, without answering from documents the asker is not allowed to see, is the half that takes the time — and it is the half that decides whether the thing gets used in month three or quietly abandoned.
The deployment
- Model selection and sizing against a stated hardware budget rather than a wish
- Deployment on your equipment, or on infrastructure we operate — decided in writing before anything is built
- Network placement: which segment it sits on, what it is allowed to reach, and what is allowed to reach it
- No outbound inference calls — the model does not send your text to a third-party service to answer
- Logging of who asked what, because you will be asked that question by an auditor eventually
Retrieval over your own documents
- Ingestion from the file estate you already have, not a new one you have to populate
- Access control mirrored from your existing permissions — the model does not answer from a document the asker cannot open
- Every answer returns the documents it drew from, so the person reading it can go and check
- Re-indexing as documents change, so the answer reflects the current version rather than last quarter's
- A written review of what it retrieves well and what it retrieves badly, before anyone relies on it
Business process automation
- Integration between systems you already own — quoting, accounting, scheduling, document management
- Elimination of manual re-entry, which is where the errors and the wasted hours both live
- Scheduled reporting that arrives without someone assembling it by hand every Monday
- Document extraction and routing where the input is structured enough to be reliable
- Priced at $165/hour with an estimated hours range per workflow
Keeping it running
- Operating system, runtime and model updates on a defined cadence
- Capacity and performance monitoring, with alerts on defined SLAs
- Backup of the index and the configuration, with restores tested
- Periodic review of what people are actually asking and whether the answers are still good
- A documented exit: the models are open, the data is yours, and the configuration lives in a repository you can read
All Private AI Deployment work is quoted hourly at $165 per hour, with estimated hours ranges based on your specific deployment scope. There is no discovery call between you and the rate.
Who buys it
Businesses that already had this problem, before anyone called it AI.
This is not a page about how transformative language models are. The buyers who need this are the ones for whom the ordinary tools are simply unavailable: a duty of confidentiality, a regulator, or a contract clause stands between them and the chatbot everyone else has been using since 2023. Private, on-premises inference is not a trend for them. It is the answer to a problem they already have.
Say it is the #1 client need
48%
Actually earn money from it
13%
- A medical or dental practice
- You are a HIPAA covered entity. Pasting a patient record into a commercial chatbot is a disclosure to that vendor, and your Notice of Privacy Practices does not cover it. A model that runs on your own equipment removes the disclosure. It does not remove the Security Rule — see section 04.
- A law firm
- Client files carry a duty of confidentiality that predates every terms-of-service page you have ever accepted. Most firms have resolved this by forbidding the tools outright, which means the associates use them anyway on personal accounts. A system inside your own network is the version of that decision you can actually supervise.
- A customs broker or freight forwarder
- Entry summaries, commercial invoices and client trade data are competitively sensitive and, under the CTPAT Minimum Security Criteria, subject to controls you have committed to in writing. Your operations run on documents. Searching them without shipping them to a third party is the requirement.
- A defense subcontractor
- Controlled unclassified information under DFARS 252.204-7012 cannot go into a commercial service that has not been assessed for it. On-premises is the only shape this can take — and there is a hard limit on our side of it, stated plainly in section 04.
The proof, and its limit
We run self-hosted inference. Your data stays yours.
This service exists because we live with the consequences of running and securing self-hosted infrastructure every day. The systems we build for clients use the same operational standards, security practices, and monitoring we apply to our own production environment. When clients need assurance that a deployment will remain available and secure, we can point to operational evidence — not promises.
Here is the limit on that claim, stated plainly: keeping your data on-premises is not the same as being compliant. The data does not leave your control, which is the whole point and a real, substantial thing. It is also the only thing self-hosting does. It does not by itself satisfy HIPAA, CTPAT, DFARS, or any other standard. The Security Rule, audit logging, access control, and encryption still apply to data on your own hardware. We do not hold AI certifications, and we will not claim compliance based on where the server is plugged in.
Out of scope
Five things this does not do. The first two are the reason to hire us.
Every one of these is something a competitor’s demo will imply the opposite of. Read them as the specification rather than as a disclaimer: a provider who will not tell you where a system fails has not thought about where it fails.
- It will be wrong sometimes, and we will not tell you otherwise
- A retrieval system searches your documents and a language model writes an answer from what it found. Both halves fail. The search can miss the right file or rank the wrong passage; the model can summarise a passage it misread, fluently and with no signal that it has done so. We will not quote you an accuracy figure, we will not tell you it does not hallucinate, and we would be sceptical of anyone who does. What we build in instead is verifiability: every answer names the documents behind it so a person can open them, and the system is scoped so that a person is always the one who decides.
- Self-hosting is not compliance, and anyone who says otherwise is selling you an audit finding
- Running the model on your own hardware means your documents do not leave your control. That is the whole point, and it is a real and substantial thing. It is also the only thing it does. It does not by itself satisfy the HIPAA Security Rule, the CTPAT Minimum Security Criteria, or DFARS 252.204-7012 and NIST SP 800-171. Protected health information put into a self-hosted system is still protected health information: access control, audit logging, encryption, workforce training and your risk analysis all apply to it exactly as they apply to your practice-management system, and where we operate the system a Business Associate Agreement is required. Those controls are work, they are quotable, and they are not a side effect of where the box is plugged in.
- We will not host controlled unclassified information
- CUI on infrastructure we operate would make us a cloud service provider owing FedRAMP Moderate equivalency and an annual third-party assessment with no permitted open findings. That is out of reach at this size and we are not going to pretend otherwise. A defense-adjacent deployment is built inside your own enclave, on your hardware, and the support arrangement is written to keep CUI out of our infrastructure rather than to quietly tolerate it arriving there.
- No unattended workflow, and no output that nobody reads
- We will not build a system where a generated answer goes into a filing, a diagnosis, a bid, a customs entry or a legal position without a person reading the source it cites. That is not caution for its own sake — it is the only configuration in which the accuracy limit above is survivable. If the value case for a project depends on removing the human, it is the wrong project and we will say so at quoting time rather than after.
- Racks, circuits and cable are licensed trades
- Inference hardware is heavy and hungry, and a serious deployment often wants a dedicated circuit and somewhere to put the box. New circuits and conduit are C-10 electrical work; a rack fastened to the building and any new cabling are C-7 low-voltage contracting. California B&P §7031(b) lets a client recover every dollar paid for unlicensed contracting work, so we specify that work, you contract a licensed trade for it directly, and we take over at the point where it has power and a link light.
None of this makes the work less valuable. A system that finds the right three documents in four seconds instead of forty minutes, and shows you which three so you can read them yourself, is worth a great deal — and it is worth it precisely because it is not pretending to be the thing that decides.
What we need from you
Five conditions, and the first one kills most of these projects.
These appear in the statement of work rather than emerging halfway through the build. Two of them are usually uncomfortable, which is why they are here at the start.
- One question people actually ask ten times a week
- Documents someone has already organised
- An access-control answer before anything is ingested
- A hardware budget decided honestly
- A written placement decision for regulated data
Next step
Before the model, the documents. Before the documents, a look at both.
The IT Health & Risk Assessment inventories the estate, the file shares and the permissions this would have to run on top of — which is the only honest basis for quoting the build. Typically 20-40 hours at our standard rate of $165/hour. Credited in full against onboarding if you sign within 30 days.