Forge a Breast Cancer diagnostic dataset
Choose how many workup records you want and a seed. The same two numbers will always produce the same dataset — write the seed down and cite it in your methodology chapter.
Advanced: customize the model
Same directional relationships described in the methodology — strengthen, weaken, or rebalance them here without touching any code. Defaults are illustrative, anchored against general breast-cancer epidemiology literature, not derived from real clinical data.
Malignancy probability
Patient age is table-rated off the age-band base-malignancy-yield anchor (not adjustable here); these loadings multiply on top of it. Active in both stage modes.
Stage at diagnosis two-stage only
Only computed for malignant records when stage mode above is two-stage. stage_score_bands (the latent-score-to-I/II/III/IV bucketing) isn't editable here — only the score contributions feeding into it are.
Cite this dataset
This is a specific, reproducible dataset — seed, record count, stage mode, generator version, and any advanced-config changes are all part of what makes it citable. Copy a citation below, or download the full provenance manifest for a supervisor or reviewer to verify against.
This is synthetic data, not real records — safe to use freely for research, coursework, or analysis, but it should never be presented as real-world data. If this dataset (or results derived from it) appears in published, submitted, or graded work, please cite the specific seed and generator version using the citation tool below.
Also downloadable as from your account's Generation history.
Data dictionary (0 columns)
| Column | Type | Unit | Description |
|---|
Hypothesis check
Directional confirmation on this batch — sign and rough magnitude should match, exact numbers will vary run to run except when the seed is held fixed. The two Malignancy cards always show; the three Stage cards only appear when this batch was generated in two-stage mode (see the methodology page).