Pre-Training Audits
Audit training data for representation gaps, label quality, and potential sources of bias before a single epoch runs. We analyze demographic distributions, identify proxy variables, and validate annotation guidelines to prevent bias from entering the pipeline at its source.
