AWS Infrastructure for Digital Health Startups

Garik H Author

U.S. digital health startups raised $7.4 billion in the first half of 2026 alone, and rounds of $100 million or more made up 45% of all that capital, according to Rock Health’s H1 2026 funding report. Money like that buys speed. It also buys scrutiny. The same week a Series B closes, a hospital system’s procurement team sends over a security questionnaire, an enterprise legal team asks for a SOC 2 report, or a payer wants a signed Business Associate Agreement before anyone touches their data. None of that happens on the startup’s timeline. It happens on the buyer’s.

Why Healthcare Infrastructure Is Different

Digital health companies move fast because they have to. Most are still chasing product market fit with an engineering team of five to twenty people, and every week spent hardening infrastructure is a week not spent on the feature a hospital system actually wants to buy. That tension sits underneath almost every architecture decision a healthtech founder makes.

At the same time, that same small team is often responsible for some of the most sensitive personal data that exists: diagnoses, mental health notes, reproductive care history, genetic results. The regulatory weight of that data doesn’t wait for headcount to catch up. A twelve person startup running its first hospital pilot can carry the same PHI handling obligations as a 2,000 person hospital IT department, with a fraction of the staff to manage it.

Compliance also tends to arrive earlier than founders expect. It used to be a Series C conversation. Now investors routinely ask about SOC 2 status during Series A diligence, and enterprise buyers stall deals in security review rather than sign without a report in hand. Getting AWS infrastructure for healthtech startups right in year one is a lot cheaper than rebuilding it under deal pressure in year three.

None of that is an argument for building enterprise scale you don’t have yet. The instinct to “take compliance seriously” often turns into a six account landing zone, a dedicated security operations center, and a HITRUST readiness project for a company with eleven employees and forty pilot users. That’s its own failure mode. It burns the exact runway the small, fast team was supposed to protect.

What Kind of Health Data Are We Actually Talking About?

Not Every Digital-Health Startup Handles the Same Data

“Healthtech” covers wildly different data realities, and the architecture that fits one profile can be badly wrong for another.

Company profile Data it actually touches Regulated as PHI under HIPAA?
Telehealth platform contracted with a physician group Video visit recordings, clinical notes, e-prescriptions Yes, as a business associate of the physician group
Remote patient monitoring startup streaming vitals to a hospital care team Device telemetry tied to a named patient, care plan notes Yes, the moment telemetry is linked to an identified patient and a covered entity
Insurance eligibility or benefits navigation platform integrated with a payer Eligibility data, plan details, sometimes diagnosis codes for prior authorization Yes, as a business associate of the payer
D2C mental wellness app with journaling and mood tracking, no clinical relationship Self-reported mood logs, journal entries Not HIPAA, but very likely “consumer health data” under state laws like Washington’s My Health My Data Act
Fertility or period tracking app Cycle data, pregnancy status, location at time of logging Same as above: outside HIPAA, squarely inside newer state consumer health data law
Claims analytics company working from de-identified data under a data use agreement Claims data stripped of the 18 HIPAA identifiers Not PHI if de-identification meets Safe Harbor or Expert Determination, but proving that holds is entirely on the company

That insurance eligibility row isn’t hypothetical. In February 2026, TriZetto Provider Solutions, a claims processing vendor rather than a hospital, disclosed a breach dating back to November 2024 that ultimately affected more than 3.4 million people tied to insurance eligibility verification records. It was the largest healthcare breach reported to HHS through the first half of the year. The company that got breached was a business associate, exactly the profile of a growth-stage digital health vendor, not a covered entity with a hospital’s compliance budget.

Identify What Is Actually Sensitive

Sensitivity isn’t a fixed label attached to a data type. It’s contextual, and it depends on identifiability and linkage as much as content.

A resting heart rate reading from a fitness wearable, on its own, is nearly meaningless. The same reading, tied to a named user with a diagnosed arrhythmia, feeding a clinician’s monitoring dashboard, is data that could change a treatment decision if it leaked or if it was tampered with. Same number, completely different risk profile, because of what it’s attached to and who can act on it.

This is where a genuinely counterintuitive trap shows up. Plenty of founders assume that if they’re not a hospital, insurer, or clearinghouse, HIPAA simply doesn’t apply, and they relax accordingly. HIPAA does not regulate health information created or held by a company with no covered entity relationship, which is true. But Washington’s My Health My Data Act, in effect since March 2024, was written specifically to close that gap. It covers “consumer health data” collected by apps, wearables, and websites that HIPAA never touches, applies regardless of company size or revenue, and carries a private right of action, meaning individuals can sue directly rather than wait for a regulator. A period tracker, a symptom checker, or a wellness app with no clinical relationship at all can be squarely in scope even while being completely outside HIPAA. “We’re not handling PHI” is not the same statement as “we have no sensitive data compliance exposure,” and the two get conflated constantly.

Data Classification Should Drive Architecture

The question founders usually start with is “Which AWS service should we use?” That’s the wrong starting point. The right one is: what data are we storing, who can access it, and why?

Answer that, and the architecture mostly falls out of the answer rather than the other way around. The flow looks like this:

Data → Classification → Access → Storage → Encryption → Monitoring → Retention

Data: what you actually collect, field by field, not what your privacy policy says in the abstract.
Classification: PHI, other PII, de-identified, or public. Each tier gets different handling, not the same S3 bucket with different folder names.
Access: who and what, human or service, can reach that data, expressed as IAM policy rather than a shared assumption that “the backend team has access to everything.”
Storage: where it physically lives. PHI belongs in an isolated, tightly scoped account or VPC, not comingled with marketing analytics or application logs.
Encryption: at rest with a customer-managed KMS key you control the rotation and access policy for, in transit with TLS 1.2 or higher, everywhere, no exceptions for internal traffic.
Monitoring: who touched what, and when, captured automatically through CloudTrail and reviewed, not just logged and forgotten.
Retention: how long you’re actually required to keep it, often six to ten years depending on state medical records law, and not a day longer than that requirement plus your own operational need.

Skip step one and every later decision compounds the mistake. Teams that pick “least privilege IAM” or “encrypt everything” as a starting principle without first classifying the data end up either over-restricting a low-risk analytics pipeline into uselessness or under-restricting the one production database that actually holds PHI.

Data Minimization

The cheapest control in this entire article is not collecting data you don’t need. Every field you don’t store is a field that can’t leak, doesn’t need a retention policy, doesn’t need to be de-identified before it reaches an analytics warehouse, and doesn’t inflate the record count if there’s ever a breach to report. IBM’s 2026 Cost of a Data Breach Report ties breach cost directly to the volume and sensitivity of records exposed, so a smaller PHI footprint is a smaller number on that curve before anything else changes.

In practice this looks like storing a hashed patient reference instead of a full identifier where a downstream service only needs to match records, truncating free-text fields that don’t need full clinical narrative, and keeping PHI in one segmented data store so a compromise of the marketing stack or the support ticketing tool never touches it at all.

Where Compliance Actually Lives: The Technical Controls

This isn’t a services catalog, and a full walkthrough of every control belongs in its own article. But at a glance, here’s where the real risk concentrates and what AWS gives you to close it.

Area What actually goes wrong The AWS-native answer
IAM Overly broad roles, shared service accounts, no periodic access review Role-based policies scoped per service, IAM Access Analyzer to flag unused permissions, IAM Identity Center for human access instead of long-lived keys
Network isolation PHI workloads sitting in the same VPC and security groups as public-facing marketing infrastructure Dedicated VPCs or accounts per environment, private subnets for anything touching PHI, VPC endpoints instead of routing through the public internet
Encryption & secrets Default AWS-managed keys with no audit trail of who can decrypt, credentials hardcoded in application config Customer-managed KMS keys with key policies and CloudTrail-logged usage, Secrets Manager with automatic rotation instead of environment variables
Logging & auditability CloudTrail enabled per account instead of organization-wide, log buckets that can be tampered with by the same admins being audited Organization trail delivering to a locked-down log archive account, S3 Object Lock on the log bucket so nobody, including root, can delete history
Sensitive information in logs Full request bodies or patient identifiers landing in application or access logs by default Field-level redaction before logging, Macie scanning S3 for accidental PHI exposure, log retention policies that don’t quietly become a second unmanaged PHI store
Backups & disaster recovery Backups that exist but were never tested, or that live in the same region and account as the production data they’re supposed to protect AWS Backup with cross-region copy, a documented and periodically tested restore procedure, not just a green checkmark on a backup job
IaC & change management Manual console changes that drift from what Terraform or CloudFormation says the environment should look like Infrastructure as code as the only path to production, drift detection, and Service Control Policies that block the console changes people make under deadline pressure

Every one of those seven areas has specific controls, specific AWS Config rules, and specific traps that show up in real audits, and listing all of them here would turn this article into exactly the services catalog it’s trying to avoid being. We’ve already written that checklist end to end, control by control, including the IaC drift and Service Control Policy tamper guards that most teams miss on their first pass, in “HIPAA AWS Checklist: Controls You Need Before You Go Live.” If you’re trying to verify your own setup against a real list rather than a vibe, that’s the piece to read next.

It’s worth saying HIPAA and SOC 2 Are Not the Same, because the two get treated as interchangeable constantly: HIPAA is a legal requirement tied to specific data and specific relationships, and SOC 2 is a voluntary attestation that enterprise buyers ask for as a proxy for operational maturity. You can be fully HIPAA compliant and still lose an enterprise deal because you don’t have a SOC 2 report, and you can hold a SOC 2 report and still be in violation of HIPAA if you’re mishandling PHI the audit never looked at.

The two frameworks overlap more than most teams assume, roughly 65% of controls by Vanta’s estimate, which is exactly why so many startups try to run both audits off one control set and end up with gaps in the 35% that doesn’t overlap. We cover where that gap actually sits, and how to sequence the two audits instead of running them blind, in “HIPAA vs SOC 2: Do You Need Both?”

A Realistic AWS Architecture for a Digital-Health Startup

Picture Pathlight Health, a fictional composite built from patterns we see repeatedly: a 60 person, Series B remote patient monitoring company selling into three regional hospital systems, with a signed BAA in place and its first SOC 2 Type II audit six months out. This is the profile the rest of this section is built for, not a five person pre-seed team and not a 500 person public company.

At this stage, a single flat AWS account with everything in it is no longer defensible, but a fifteen account enterprise landing zone is overkill. A right-sized foundation looks like this:

Organization (AWS Organizations)

AWS Healthcare Cloud Architecture

This mirrors the account separation AWS itself recommends in its Landing Zone Accelerator for Healthcare, which uses a comparable Security OU, Infrastructure OU, and a dedicated OU for PHI-bearing workloads across roughly 35 AWS services, deployed through Control Tower and CloudFormation (AWS Public Sector Blog, June 2024). The blog post is explicit about one thing worth repeating here: the landing zone gives you the scaffolding, it does not make you compliant by itself. Pathlight’s version above is a scaled-down variant, three to five accounts instead of the full enterprise OU tree, because a 60 person company doesn’t yet have the workload diversity that justifies more.

Inside the production account, the components that actually carry PHI and the traffic around it:

Component Role Why it matters at this stage
CloudFront + AWS WAF Edge termination, managed rule groups against common web exploits Stops the noisy, automated attacks before they ever reach compute, cheaply
ALB → ECS Fargate (or EKS) Application compute Fargate removes host patching from a five person platform team’s plate entirely
Aurora PostgreSQL, Multi-AZ, KMS-encrypted System of record for structured PHI Automated failover matters when a hospital’s care team is depending on real-time vitals
Amazon S3, encrypted, access-logged Documents, attachments, device exports Object-level access logging gives you the audit trail regulators and auditors actually ask for
Amazon Cognito or an equivalent managed IdP User and clinician authentication MFA and session policy enforced centrally instead of reimplemented per service
Secrets Manager Database credentials, third-party API keys Automatic rotation without a deploy, and nothing sensitive sitting in an environment variable
GuardDuty + Security Hub + Config, org-wide Threat detection, standardized findings, continuous configuration compliance One place to look during an incident instead of four separate consoles
AWS Backup with cross-region copy Disaster recovery A second region holds the restore point, not just a second folder in the same bucket

What This Actually Costs

A basic, non-healthcare startup running a simple web app on EC2, RDS, S3, and CloudFront typically lands in the $200 to $800 a month range before traffic scales meaningfully. Add a real HIPAA-aligned compliance posture, multi-account separation, organization-wide security tooling, and cross-region backup, and the number moves into a different bracket entirely. For a company at Pathlight’s stage, a reasonable range looks like this, infrastructure only, no headcount:

Category What’s in it Typical monthly range
Compute & containers ECS Fargate or EKS, load balancer, autoscaling $800 to $2,500
Database Aurora PostgreSQL Multi-AZ, encrypted, one read replica $600 to $2,000
Storage & CDN S3, CloudFront, EFS if needed $200 to $600
Security & compliance tooling GuardDuty, Security Hub, Config rules across all accounts, Macie on PHI buckets $500 to $1,500
Logging & observability Org-wide CloudTrail, CloudWatch Logs with real retention, alerting $300 to $900
Backup & disaster recovery AWS Backup, cross-region copy, tested restore $300 to $1,000
Network security WAF, NAT gateways, VPC endpoints, Shield Standard $200 to $700
AWS Support plan Business Support, effectively required once PHI is in production ~$100 minimum or roughly 3% of spend

That lands somewhere around $2,900 to $9,200 a month in pure AWS spend, before a single engineer’s salary, and it moves in either direction with region, traffic, and how aggressively the team uses Savings Plans versus on-demand pricing. It’s a real, board-visible line item, which is exactly why so many growth-stage teams end up buying managed AWS for digital health from a partner rather than hiring the two or three additional platform and security engineers it would take to run all of this internally at the same standard.

Match the Architecture to the Company You Actually Are

If you’re a pre-seed team, a growth-stage company, or a mature enterprise, what “right” looks like is genuinely different, and copying the wrong one is one of the most expensive mistakes a healthtech team can make in either direction.

Pre-seed and MVP teams, generally under 15 people, are usually running on a handful of design partners with PHI volume near zero or fully synthetic. One AWS account is right-sized here: managed services like Aurora Serverless, Fargate, and Cognito, a signed BAA, encryption on by default, and disciplined use of HIPAA-eligible services. A multi-account landing zone, a dedicated security hire, or HITRUST readiness at this stage is overkill, not diligence.

Growth-stage companies, in the 25 to 100 person range, are the ones with signed BAAs with real health systems, a SOC 2 report requested mid-sales-cycle, and their first external security review on the calendar. That’s exactly the reality the account structure and controls in this article are built for, along with a named security owner even if that role is only fractional. A fifteen-account enterprise landing zone, a 24/7 in-house SOC, or a custom-built SIEM are not proportionate yet.

Mature or enterprise companies, 500 or more people, running multiple product lines, sometimes a covered entity in their own right, with a global footprint, are the ones that actually justify dedicated platform and security teams, HITRUST or ISO 27001 certification, multi-region active-active infrastructure, and a formal change advisory board. The overkill risk flips at that scale: it’s skipping systemic controls in favor of one-off fixes for whatever the last audit happened to flag.

Don’t build that mature architecture on day one just because you’re a healthcare startup. Build the one that matches the data you actually hold, the deals you’re actually closing, and the team you actually have, and let it grow account by account as those things change. That’s a harder discipline than defaulting to “more security is always better,” but it’s the one that keeps both the compliance posture and the runway intact.