Forge a health insurance dataset
Choose how many policy records you want and a seed. The same two numbers will always produce the same dataset — write the seed down and cite it in your methodology chapter.
Advanced: customize the model
Same directional relationships described in the methodology — strengthen, weaken, or rebalance them here without touching any code. Defaults are illustrative, not derived from real health-utilization data.
Claim cost (severity)
Claim frequency — utilization loadings
Age is table-rated off the age-band claim-incidence anchor (not adjustable here); these loadings multiply on top of it. Raise or lower them to strengthen or weaken the chronic-condition/smoker/ BMI hypotheses.
Name field
Off by default — most statistical work has no use for a name column. This is a separate decision from the Country field above: that only decides which name pool a name would be drawn from, not whether this dataset has names in it at all.
Cite this dataset
This is a specific, reproducible dataset — seed, record count, generator version, and any advanced-config changes are all part of what makes it citable. Copy a citation below, or download the full provenance manifest for a supervisor or reviewer to verify against.
This is synthetic data, not real records — safe to use freely for research, coursework, or analysis, but it should never be presented as real-world data. If this dataset (or results derived from it) appears in published, submitted, or graded work, please cite the specific seed and generator version using the citation tool below.
Also downloadable as from your account's Generation history.
Data dictionary (0 columns)
| Column | Type | Unit | Description |
|---|
Hypothesis check
Directional confirmation on this batch — sign and rough magnitude should match, exact numbers will vary run to run except when the seed is held fixed. Like life, these hypotheses are on claim frequency, not claim cost (see the methodology page).