Synthetic WH-380-F Certification of Health Care Provider for Family Member's Serious Health Condition Data

Synthetic training data — no real PII, fully coherent identities

HR2024

Generate synthetic WH-380-F FMLA medical certifications for a family member's serious health condition, with a named patient, a stated care relationship, and a provider specialty that fits the diagnosis. Structurally the sparsest form in the FMLA family and the classic adversarial partner to the WH-380-E.

26

Fields per document

4

Pages

HR

Category

What this document is

Form WH-380-F is the U.S. Department of Labor's model certification of a health care provider for a family member's serious health condition — the certification an employee files when the leave is to care for a spouse, child or parent rather than for themselves. It identifies the employee, the family member being cared for, and the patient's treating provider, then asks the provider to describe the condition, its expected duration, and the nature and frequency of the care the employee will need to provide.

Why generate synthetically

Caregiver leave is the fastest-growing category of FMLA usage and the hardest to source documents for: the medical content is about someone who is not the employee, which makes an already sensitive record sensitive on behalf of a third party. Real WH-380-F files essentially never leave the administrator that holds them. Synthetic certifications give leave-management vendors a caregiver corpus of arbitrary size, and — more importantly for model quality — give them the E-versus-F contrast set that a corpus assembled from one employer's files almost never contains in balance.

What makes synthetic data useful

Each certification names a real relationship inside the simulated household. The family member being cared for is drawn from the employee's own simulated family, the patient name and the family-member name agree, and the care description follows from the diagnosis: a chemotherapy course produces assistance with daily activities and transport to treatment; a post-stroke rehabilitation produces a different care pattern and a different provider specialty. Condition duration, incapacity frequency and incapacity duration are drawn as a set so the timeline the administrator would compute from the form actually resolves.

Training challenges

This is a four-page document with very few fields on it, which is its own extraction problem: page after page of dense printed regulatory instruction with two or three answer regions floating in it, and a model tuned on dense forms will over-predict. The employee name repeats as a header on every page while the patient name appears once, and confusing the two sends a caregiver certification into a self-certification workflow. Almost every answer is free text of unpredictable length rather than a coded field, so there is little geometric regularity to lean on, and the single checkbox on the form — whether the condition is chronic — is the only structured clinical answer in the entire document.

Generate synthetic WH-380-F Certification of Health Care Provider for Family Member's Serious Health Condition data

Start with 500 free credits. No credit card required.

Generate Now

Who uses this data

FMLA and absence-management administrators and the third-party administrators who process caregiver leave at scale, leave-management platforms building certification intake, and disability and paid-family-leave carriers whose intake forms share this structure. It is also the sharpest available negative example for anyone building an FMLA document classifier around the WH-380-E.

Document complexity profile

26 fields across four pages — 25 text answers and a single checkbox — with 29 annotation relations. That is the lowest field density of any form in the catalog whose fields span every page: four pages of printed regulatory text carrying a handful of free-text answers. Binding logic is light (3 conditional bindings, 9 function calls, maximum expression depth 2); the clinical coherence comes from a registered computed-field module rather than from expressions on the page.

Key stats from our synthetic corpus

Quantitative characteristics of the WH-380-F Certification of Health Care Provider for Family Member's Serious Health Condition documents our generator produces.

MetricValueDetail
Chronic condition split51% / 49%The one clinical checkbox on the form marks the condition chronic on 50.7% of certifications. A near-even split is deliberate: with a single structured binary on the whole document, an imbalanced corpus would let a model score well by ignoring it entirely.
Field density26 fields / 4 pagesFewer than seven answer regions per page on average, with the remaining area given over to printed instructions. Sparse documents are where dense-form extractors over-predict, and this is the sparsest form in the catalog whose fields span every page.
Provider specialties represented10Certifying providers span ten specialties, led by family medicine at 32% and internal medicine at 24%, with oncology, cardiology, neurology and others in the tail. Specialty is paired to the patient's condition rather than sampled independently.
Distinct care scenarios4Four caregiver scenarios circulate — active chemotherapy, post-stroke rehabilitation, post-surgical recovery, and long-term condition management — each with its own condition duration, incapacity frequency, and care description. Duration answers range from a six-to-eight-week recovery to a permanent condition.
Employee name repeats per document4The employee's name appears in a header field on all four pages while the patient's name appears once. Multi-page documents whose pages carry a repeated identifier are the standard test for page-grouping and out-of-order-scan handling.

How this document co-occurs with others

Rates at which identities in our corpus that produce a WH-380-F Certification of Health Care Provider for Family Member's Serious Health Condition also produce other documents.

CorrelationRateDetail
Self-certification counterpart100%The same employee can file a WH-380-E for their own condition. Generating both gives you the caregiver-versus-self contrast pair over a shared layout family — the training set for the classification decision that determines which FMLA workflow the document enters.
Employer's eligibility notice for the same leave100%The employer's WH-381 notice of eligibility accompanies the certification in the leave file, and its family-relationship checkboxes should agree with the family member named here. Disagreement between the two is a real compliance finding, and the pair supervises detecting it.
Beneficiary designation naming the same family100%The employee's benefits beneficiary designation names the same simulated spouse, children and parents that supply the patient here. Cross-document family-relationship consistency is a signal HR platforms use and almost never have labelled data for.
Employees who also file as seniors2%2.3% of FMLA-eligible identities also qualify for Form 1040-SR. Caregiver leave skews toward employees with ageing parents, and the small overlap with senior filers is a reminder that the caregiver and the patient sit in different age bands by construction.

Rate figures above are corpus-derived: they were computed over the 641 employment-eligible identities inside a local synthetic corpus of 1,000 identities generated by SymageDocs' World Simulation Engine at seed 20260421. Field, page, type and relation counts come from the shipped WH-380-F definition in the SymageDocs form library. No real employee, patient, or provider data was used at any stage.

Frequently asked questions

What data format do synthetic WH-380-F documents include?
Each generated identity produces a rendered PDF plus a structured JSON annotation file with bounding boxes, field types, and ground-truth values for all 26 fields across the four pages, plus 29 label-to-value relations tying the printed questions to their answers. COCO, YOLO, FUNSD, and BIO/NER exports come from the same job.
Is the family member a real part of the simulated household?
Yes. The patient is drawn from the employee's own simulated family rather than invented independently, so the name that appears in the family-member block matches a person who exists elsewhere in that identity's document set. That is what makes the corpus usable for the cross-document relationship checks a leave administrator performs when a certification arrives.
Why does a form with so few fields need four pages?
Because the Department of Labor's model form devotes most of its length to printed instructions and regulatory notices. That ratio is the point: a sparse document dominated by static text is where extraction models trained on dense tax and claim forms fail hardest, and it is a realistic profile for a large share of HR and government paperwork.
How is this different from the WH-380-E?
The WH-380-E certifies the employee's own condition and includes the employer's job-description block and an analysis of the essential functions the employee cannot perform. The WH-380-F drops those and adds the family-member identification and care-description blocks. They share a header and a visual family, which is exactly why a classifier trained on only one of them misroutes the other.
Can I use this data commercially?
Yes. Every employee, family member, provider, and clinical detail is synthetic, describes no real person or patient, and is licensed for commercial use including model training, benchmarking, and redistribution inside your own products.

Related HR Forms