FHIR Bulk Data Export: $export, NDJSON and Population Extracts
FHIR Bulk Data export lets you extract data for thousands or millions of patients in one operation, instead of making millions of individual API calls. It powers population health analytics, quality reporting, payer programs and data warehouse loads, alongside standard FHIR integration APIs. This technical guide walks through the Bulk Data Access Implementation Guide: the kick-off request, asynchronous polling, NDJSON output, group versus system export and backend authorization. It also explains why naive downstream processing fails at population scale. Planning a bulk pipeline? Talk to our FHIR data engineers.
How FHIR Bulk Data Export Works
The FHIR $export operation follows an asynchronous request pattern, because population-level exports can take minutes or hours to prepare. The client sends a kick-off request, the server accepts it and returns a status URL, and the client polls that URL until files are ready. The server then returns a manifest listing NDJSON files by resource type. This approach, sometimes called flat FHIR, avoids timeouts and lets servers generate large exports without blocking other API traffic.
The Kick-Off Request
The client calls $export at the system, patient or group level, sending Accept application/fhir+json and Prefer respond-async headers. Optional parameters like _type and _since limit which resources and dates are included.
The 202 Accepted Response
The server responds with HTTP 202 Accepted and a Content-Location header containing the status URL. The client stores that URL, because every later step in the export depends on it.
Polling for Status
The client polls the status URL. While processing continues, the server returns 202, often with an X-Progress header. Clients should respect Retry-After headers and avoid aggressive polling that triggers rate limits.
The Completion Manifest
When the export finishes, the server returns 200 with a JSON manifest listing output files, each with a resource type and download URL. Error files, if any, appear in a separate list.
Downloading and Cleanup
The client downloads each NDJSON file, often requiring authorization, then sends a DELETE to the status URL when finished. This lets the server clean up files and free storage promptly.
Group, Patient and System-Level Export
The Bulk Data specification supports three export levels, and choosing the right one affects performance, authorization and data scope. Payers and accountable care organizations usually export defined member groups, while analytics teams may need broader data from sources covered by EHR/EMR integration. Many EHRs support only some levels, and each level can require different permissions. Confirm what each server supports before designing your pipeline, because switching levels later often changes authorization, scheduling and downstream data models.
Group-Level Export
Group-level export extracts data for a defined cohort, such as an attributed patient panel, using Group/[id]/$export. It is the most common pattern for payers, ACOs and value-based care programs. Group definitions must stay current.
Patient-Level Export
Patient/$export extracts data for all patients the client is authorized to access. It suits organizations with broad data rights, but can produce very large files that require careful downstream processing.
System-Level Export
System-level $export extracts all data on the server, including non-patient resources. It is usually limited to internal analytics or migration use cases, because few external clients receive this level of access.
Incremental Exports With _since
The _since parameter returns only resources changed after a given time. Incremental exports dramatically reduce file sizes and processing time, making daily or weekly refreshes practical for large populations. Track export timestamps carefully.
Filtering With _type
The _type parameter limits exports to specific resources, such as Patient, Observation and Condition. Requesting only needed resources shortens export time, reduces storage costs and simplifies downstream transformation work considerably.
Authorization for Bulk FHIR Exports
Bulk FHIR analytics workflows run without a user present, so they use SMART Backend Services authorization rather than interactive login. The client authenticates with a signed JSON Web Token, using an asymmetric key registered with the server, and receives a short-lived access token with system-level scopes. Because these exports move large volumes of protected health information, key management, scope design and audit logging deserve extra attention, as the HIPAA technical safeguards require.
Registering the Client
The client registers its public key, often as a JWKS URL, with the FHIR server. The private key stays in a secure vault and is never shared, logged or stored in code.
Signed JWT Assertions
To request a token, the client signs a JWT with its private key and posts it to the token endpoint. The server validates the signature and issues a short-lived access token.
System Scopes
Bulk exports use system scopes, such as read access to specific resource types. Request only the resources your pipeline needs, because broad scopes increase risk and slow down approval. Document every scope requested.
Key Rotation
Rotate keys on a defined schedule and whenever staff with access leave. Publishing keys through a JWKS URL makes rotation easier, because servers can fetch new keys without manual updates.
Auditing Export Activity
Log each export request, its scope, files downloaded and destination systems. These records support HIPAA audit controls and help answer questions about who accessed population data and when. Retain logs per policy.
Processing NDJSON at Population Scale
The export itself is often the easy part. Problems appear downstream, when pipelines try to load gigabytes of NDJSON into databases or analytics tools. Loading entire files into memory, processing records one at a time or ignoring references between resources leads to failures, slow jobs and inconsistent data. Well-designed pipelines stream files, validate records and load data in batches. We build these pipelines as part of our healthcare API development and data engineering work.
Stream, Do Not Load
NDJSON stores one resource per line, so pipelines should stream files line by line. Loading whole files into memory fails quickly once exports reach millions of records per resource type.
Batch Database Writes
Write records in batches rather than individually. Batched inserts or bulk load tools reduce database overhead dramatically and turn multi-hour loads into jobs that finish in minutes. Tune batch sizes through testing.
Resolving References
Observations reference Patients and Encounters by ID. Load parent resources first, or use staging tables, so references resolve correctly and analytics queries join data reliably. Orphaned references are logged and investigated, never silently dropped.
Validation and Error Handling
Validate resources against expected profiles and log failures separately. One malformed record should never stop an entire load, but repeated failures should alert your data team quickly. Error files need review.
De-identification for Analytics
Many analytics use cases do not need identifiable data. Our healthcare data anonymization service de-identifies exported data before it reaches research or AI environments, reducing risk significantly. Identifiers stay in secured environments only.
Frequently Asked Questions
What is FHIR Bulk Data export?
FHIR Bulk Data export is an asynchronous operation, called $export, that extracts large volumes of FHIR resources for many patients at once. The server prepares NDJSON files and returns a manifest of download links. It is defined in the HL7 Bulk Data Access Implementation Guide and is widely used for analytics, quality reporting and payer programs.
What is NDJSON in FHIR?
NDJSON, or newline-delimited JSON, stores one complete JSON resource per line. FHIR Bulk Data uses it because files can be streamed and processed line by line, without loading everything into memory. Each file usually contains one resource type, such as Patient or Observation, making downstream processing simpler.
What is the difference between group and system export?
Group export extracts data only for patients in a defined group, such as an attributed panel, which suits payers and ACOs. System export extracts all data on the server, including non-patient resources. System-level access is rarely granted to external clients and is mostly used for internal analytics or migrations.
How do you authorize FHIR Bulk Data requests?
Bulk Data uses SMART Backend Services authorization. The client registers a public key with the server, signs a JWT with its private key, and exchanges it for a short-lived access token with system scopes. No user login is involved, so secure key storage and rotation are essential.
How long does a FHIR Bulk Data export take?
Export time depends on population size, resource types, server load and vendor implementation. Small groups may finish in minutes, while large populations can take hours. Using _type to limit resources and _since for incremental exports reduces time considerably, making regular scheduled refreshes practical for most organizations.
Do all EHRs support FHIR Bulk Data export?
Many certified EHRs support Bulk Data export, especially group-level export, because it is part of US certification requirements. However, supported levels, resource types, parameters and performance vary by vendor and customer configuration. Always confirm capabilities with each server and test against real environments before committing to production timelines.
Planning a FHIR Bulk Data Pipeline?
Our integration engineers are ready to help. Free consultation, no obligation.