Both packages consume the same four kinds of input. This section documents each one’s schema, shows worked examples drawn from the bundled demo dataset, and exposes every artefact as a download — so you can preview the shape, grab the standalone CSV/TSV files, or load the canonical .rds directly with load_data().
Each subpage uses the same canonical dataset, shipped byte-identically by both sibling packages. Two scales are bundled:
Size
Individuals
SNPs
Microhaplotype blocks
\(W_{ah}\) columns
QTL
Use
large (main)
200 (10 sires × 20 offspring)
500
250
~735
20
All standard examples
small (toy)
100 (10 sires × 10 offspring)
50
25
~75
5
Quick-start (runs in seconds)
Both are simulated with set.seed(42) for the genetic stream and set.seed(41) for the within-family train/test split; re-running demo-data/generate.R produces bit-identical output.
TipOne .rds per size, byte-identical to the sibling packages
demo-data/main/demo_data.rds and demo-data/toy/demo_data_small.rds are byte-identical to the files shipped under masbayes/inst/extdata/ and masreml/inst/extdata/. The canonical generator lives in the workspace as tools/make_demo_data.R; the local mirror is demo-data/generate.R.
Loading
Code
# Either package's load_data() returns the canonical list.library(masreml)# or masbayesd<-load_data("large")# or "small"# Equivalent via system.file() (returns the bundled path):path<-system.file("extdata", "demo_data.rds", package ="masreml")d<-readRDS(path)# Or load the standalone .rds directly from the site repo:d<-readRDS("demo-data/main/demo_data.rds")
d is a named list with 13 slots. Twelve carry the single-generation demo (snp, mh, allele_freq, pheno, pedigree, qtl, meta, family_id, train_idx, test_idx, map_snp, map_mh) — the four subpages describe each schema and common ways to construct it from your own data. The thirteenth slot, multigen, holds a 3-generation forward-prediction extension (gen1, gen2_snp, gen2_mh, gen3_snp, gen3_mh, three reference_structure_* objects, plus an OCS log) used by tutorials/forward-prediction.qmd. Single-generation tasks are unaffected: legacy code that touches d$snp etc. ignores the nested d$multigen entirely.
Within-family train/test split
The bundle ships a stratified split: each of the 10 families contributes 16 offspring to training and 4 to test in large (8/2 in small). This keeps full-sibs out of the test partition’s predictor pool without losing family-level relatedness.
Size
n_train
n_test
per family
large
160
40
16 / 4
small
80
20
8 / 2
Indices live in d$train_idx and d$test_idx (row indices into d$snp, d$mh, d$pheno).
The ZIPs bundle the canonical .rds plus every derived CSV/TSV used on the subpages.
Regenerating locally
demo-data/generate.R is self-contained — it builds both the single-generation layer and the multi-generation $multigen extension end-to-end in one pass. Requires masbayes (uses construct_wah_matrix()), AlphaSimR, and optiSel:
Rscript demo-data/generate.R main toy # both sizesRscript demo-data/generate.R main # large onlyRscript demo-data/generate.R toy # small only
Re-running emits the same .rds byte-for-byte (set.seed() at the top of every config). The combined dual-layer file matches what the sibling packages ship under inst/extdata/, so load_data() from either masbayes or masreml returns the same list.