The best training data you could want is the run that already happened in production — a real ticket, an agent's pull request, a reviewer's two changes, a clean merge, three weeks live with no rollback — and at most companies it gets deleted at midnight. Maya and Leo follow that record out of a live coding-agent product and ask how its exhaust becomes usable training data. Telemetry hands you four labels reality wrote for free (did it merge, did the developer rewrite it, did the comments get applied, did it survive post-merge), but it is guilty until proven clean: it has to run a gauntlet of four gates — privacy, secret-scan, license, contamination — and fail any one and it is not data. Then they take opposite sides of the real split: manufacture clean synthetic tasks you fully control, or mine the radioactive-but-real production stream. The resolution is a division of labor — synthetic for volume and coverage, telemetry for distribution-truth and reality's labels — and the quiet punchline that the teams who get to learn from production are the ones who built the governance plumbing first.
Imagine the perfect training example: a real engineer's ticket, a coding agent's pull request, a human reviewer's two requested changes, a clean merge, and three weeks in production with no rollback — every label you would pay for, written by reality for free. At most companies that record is deleted at midnight because it ran in a customer's private repo. Maya and Leo use that to unpack module 5.4 — turning production telemetry into training data. Telemetry is the exhaust of a live product (branches, reviews, CI runs, merges, post-merge outcomes), and its appeal is simple: a synthetic task is something you describe, a telemetry record is something that actually happened, carrying four reality-labels nothing else can give you. But it is radioactive, so before a single record becomes data it runs a gauntlet of four gates: privacy (whose information is in it), secret-scan (a leaked credential poisons the weights and can never be un-leaked, so the whole record is quarantined), license (are we allowed to train on this code), and contamination (does it overlap the eval set and teach the model its own exam). Then they stage the genuine split — synthetic data, clean and scalable by construction, versus production telemetry, the only real distribution — and land on a division of labor: synthetic for volume and controllable coverage, telemetry for distribution-truth and the labels reality writes for free, with the quiet punchline that whether you get to use telemetry at all is decided by whether you built the privacy, secret-scan, license, and contamination plumbing before you needed it. Governance is the permission slip, not the tax. The honest limit: even cleaned telemetry is a mirror of today's product, blind to the tasks users never tried, and its free labels are still tired-human judgments — a merge can mean 'this is good' or 'it's Friday, ship it' — so never confuse 'this happened in production' with 'this was good engineering.'
来源材料
完整单集页面