
TL;DR:
A signed BAA from your model provider covers what the provider owes you. A hospital review asks what your system does, and those are different questions.
- De-identifying the input closes one of three layers. Shipping on that basis leaves two other questions unanswered.
- The model’s own behavior needs a record: which version answered, on what context, and who was allowed to override it.
- Whether your feature is a regulated device is a separate question from whether it handles PHI correctly, and different people decide it.
- The three layers cannot be phased. A feature that satisfies one and defers two is not partially ready.
What a HIPAA-compliant LLM actually has to mean
A HIPAA-compliant LLM is a feature whose data handling, model behavior, and regulatory position can each be evidenced on request. The phrase gets used as though it described a model. It describes a system, and the model is one component inside it.
This matters because the market sells the first part and leaves the rest. Vendors will sign a business associate agreement, and a signed BAA is necessary. It establishes that the provider may handle protected health information and what it owes you if something goes wrong. It says nothing about whether your application requests the minimum necessary data, logs what the model returned, or can explain why a clinician saw a particular suggestion.
So when a founder asks whether a given tool is compliant, the honest answer starts by reframing the question. Is AI HIPAA-compliant is a question about a category, and HIPAA does not regulate categories of technology. It regulates covered entities and business associates, and it asks what they did with the data.
Corpsoft Solutions organizes this into the AI Compliance Stack: three layers that have to hold at the same time. Layer 1 governs what the model may see. Layer 2 governs what the model did and who can prove it. Layer 3 governs which rules apply to the feature at all.
Layer 1: what the model may see
The first layer covers where the input came from, on what basis, how long it persists, and in what form it reaches the model.
De-identification is the control teams reach for first, and it is genuinely useful. It is also narrower than it appears. Under 45 CFR §164.514, Safe Harbor requires removing eighteen specified identifiers and the covered entity having no actual knowledge that what remains could identify someone; Expert Determination requires a qualified person to document that the re-identification risk is very small. Neither standard is satisfied by stripping a name field.
PHI de-identification fails in predictable places once an LLM is involved. A prompt assembled from several data sources can be re-identifiable even when none of them was on its own:
| Data handling issue | Why it matters for an LLM | Engineering control |
| Free-text PHI | Identifiers sit inside the narrative, where field-level rules find nothing | Redaction applied to content, not just structured fields |
| Image or file metadata | EXIF and similar metadata can survive processing that only cleans the visible content | Metadata stripped or validated before the file reaches the model |
| Embeddings | A derived representation can preserve what redaction was meant to remove | Define whether embeddings can retain sensitive information, and apply the same access, retention, and de-identification controls to them |
| Prompt and inference logs | A stored prompt containing PHI becomes part of your PHI environment from the moment it is written | Explicit retention limits and access controls on the log store |
Retention is the other half of this layer, and it is where the quiet accumulation happens. Prompt logs and inference logs are a PHI store from the moment a patient’s symptom description enters one. Our guide AI Data Leakage: The Three Paths Sensitive Data Takes Out of Your Product traces how that data physically leaves a system; the question here is narrower — what a team decides to let in before any of that applies.
AI in healthcare compliance gets treated as this layer alone, which leaves a good answer to one question and nothing for the next two.
Layer 2: what the model did, and who can prove it
The second layer is about the model’s own conduct: which version produced an output, on what retrieved context, and what a human was able to do about it.
This layer is invisible in a demo, which is why it gets deferred. A working prototype shows that the feature produces plausible output. It does not show which model version produced it, what documents were retrieved, or whether a clinician who disagreed could record that disagreement anywhere the system would keep.
Four things belong in the record, and they are useful in the order a reviewer asks for them:
- The model and its version, captured per request. A feature that silently follows a provider’s latest model has no stable answer to what produced last quarter’s outputs.
- The retrieval context. If the output was grounded in documents, the identity of those documents is part of the answer.
- The human decision. Human-in-the-loop is a design commitment before it is a compliance term: who reviews, what they can change, and whether their override is stored.
- The authorization behind the request, tying each access to a user and a purpose.

This is where practical AI HIPAA compliance gets tested. A provider’s statement about its own security is not a substitute. Our guide GDPR-Compliant AI: What Engineers Need to Log, Trace, and Explain sets out what that record should contain; the same fields do most of the work under either regime.
Layer 3: which rules actually apply to your feature
The third layer decides which regime you are in, and teams reach it last because it looks like a legal question. It changes architecture.
Two determinations do most of the work. The first is whether your feature is a regulated device. The FDA’s Clinical Decision Support Software guidance, updated in January 2026, explains the agency’s thinking on which CDS functions fall outside the device definition under the criteria added to the FD&C Act by the 21st Century Cures Act. A feature that presents information for a clinician to consider sits in a different place from one that outputs a specific recommendation the clinician is expected to follow, and the difference is drawn in the statute rather than in your product copy.
Two outputs illustrate the line without settling it on their own: “Here are the patient’s recent lab results and the relevant guideline excerpt” sits closer to informational; “Based on these findings, prescribe X” sits closer to directive. Wording alone does not decide the classification — the statutory criteria and the software’s actual function do.
Teams searching for HIPAA FDA AI development regulations are usually circling this line without knowing it has a name.
The second determination is the HIPAA one: whether the data is identifiable, and who is acting as covered entity or business associate. Both determinations can change with a product decision as small as adding a new output field.
Questions about AI and HIPAA compliance get answered much faster once this layer is settled, because it tells you which evidence you will eventually be asked to produce.
The three layers are not a sequence

Roadmaps treat the layers as phases, and the sequence is appealing: handle the data now, add logging later, sort out the regulatory position when someone asks. Each layer is separately tractable, so the phased plan looks responsible.
It fails at the moment of examination. A hospital’s security questionnaire, an investor’s diligence list, and a regulator’s inquiry each sample across all three at once. The reviewer who is satisfied with your de-identification asks next what the model logged, and a good answer to the first question makes the silence after the second louder.
There is also an ordering trap inside the work itself. Retrofitting Layer 2 means rebuilding the message store, because logging the decision requires structure the original schema did not have. Discovering a Layer 3 determination late can invalidate a Layer 1 design, since a feature reclassified as clinical decision support inherits obligations that change what may be stored and for how long.
This is the part of HIPAA and AI in healthcare that a checklist format hides. A checklist implies items can be ticked in any order. Three simultaneous conditions behave differently: partial completion produces no partial credit.
Anyone treating HIPAA-compliance AI work as a pre-launch cleanup arrives at the same place, which is a rebuild under a deadline.
Build, buy, or route through a provider?
Three approaches are available, and the choice determines which layers you own outright.
| Approach | What you control | What you hand over | Where it leaves a gap |
| Self-hosted model | All three layers | Nothing, including the maintenance burden | Cost and expertise concentrate on your team |
| Vendor tool with a BAA | Application logic and what data you send | Layer 1 retention terms, Layer 2 logging depth | The vendor’s audit trail may not answer your reviewer’s question |
| Provider model behind your own gateway | Layer 1 and Layer 2 | Model weights and inference only | Layer 3 determination stays yours regardless |
This architecture keeps redaction, logging, and authorization inside your perimeter while letting someone else run the model. It also preserves the ability to switch providers, which matters more than teams anticipate when terms change mid-contract.
Healthcare AI software development quotes vary by which of these three a supplier assumes, and the assumption is rarely stated. Asking directly saves a round of surprises.
The middle path is worth naming for what it costs: implementing AI in healthcare through your own gateway is more engineering upfront and materially less rework later.
What moves the number before anyone quotes one
Estimates for HIPAA compliant AI development diverge for reasons that have little to do with the model and a great deal to do with scope decisions made in the first week.
Four factors move it most:
- Whether the feature touches identifiable data at all. A summarizer working on already de-identified records is a different project from one reading a live chart.
- Whether output is advisory or actionable. Writing back into a record pulls in validation, rollback, and clinical review.
- Whether you already hold a BAA covering the specific endpoint and plan. The absence of one is a procurement timeline, not an engineering one.
- Whether the Layer 3 determination has been made. An unsettled device question makes every downstream estimate provisional.
Healthcare AI development budgets are usually built around the first factor and surprised by the fourth. For a planning range before a scoping conversation, our project estimator takes a described scope and returns one. AI implementation in healthcare rarely fails on the model; it fails on the second and fourth items above, discovered late.
For teams weighing whether to build this capability internally, machine learning for healthcare describes the shape of the engagement.
What this looks like on a real system
Our HIPAA-compliant dermatology telemedicine platform is a worked example of Layer 1 and Layer 3 decisions taken before the AI feature shipped.
System. Patients upload skin images for remote assessment. An AI model performs a preliminary analysis, and a licensed dermatologist reviews the case before anything reaches the patient.
Problem. A clinical image can itself contain identifying information, and it may also arrive with metadata attached by the patient’s device.
Layer 1 decision. Images were treated as protected health information from intake onward, not run through de-identification. The choice was to keep the data identified and control access to it instead — a legitimate answer whenever de-identification cannot be made to hold.
Layer 3 evidence. The published case study maps each control to the HIPAA safeguard it addresses — the specifics are below.
Documented in the public case study:
- Encryption of images in storage
- Role-based access limited to the assigned dermatologist
- Two-factor authentication for clinician and administrative accounts
- Immutable audit logs of PHI-related actions
- A defined log retention policy
Not documented in the public case study: model version capture, retrieval logging, human-override records. Nothing here should be read as a statement about them either way.
Related questions worth exploring
Wondering what gives way once the system grows past its first production cluster? AI Data Governance: What Breaks at Scale, and What Fixes It walks through four failure points, from consent tracking to audit log volume.
Evaluating the model provider itself? Third-Party Software Security Assessment: What to Check and Why Most Teams Skip It covers how to decide which vendors get reviewed and how often.
Preparing the feature for an external regulatory review? AI Regulatory Compliance: Where Most Pilots Get Stuck — and How to Prevent Failure covers what reviewers look for when a pilot moves toward production.
Subscribe to our blog