Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution
Knowing where a generated example came from and knowing whether it is worth training on are two different questions. This experiment measures the gap between them.
Provenance as a selection rule
Repeated training on model-generated data can degrade later models. One possible response is to use provenance when deciding which generated examples to reuse. This paper tests both how reliably that provenance can be recovered and whether it helps identify better training data.
The study uses financial-risk text. It first identifies the source of generated passages, then repeats the test after rewriting them, and then compares two ways of selecting generated examples over three rounds of generation and retraining.
Attribution degrades under rewriting, and selection rules do not separate
Generator attribution is 98.7% accurate on the original passages. It falls to 53.1% after paraphrasing and to 29.0% after style rewriting. Generated-versus-human detection remains close to perfect against the tested human comparison set.
Of the two selection rules, one uses source information and the other uses a score from a separate reference model. The two rules select different examples, but the planned comparison does not detect a stable difference in the degradation of the resulting models.
The results show that identifying where data came from and identifying which data are useful for training are separate problems. The experiment separates source identity, criterion-facing selection, and recursive training outcome. Neither the provenance score nor the tested criterion-facing proxy is established as sufficient for future recursive behaviour.
What the paper does and does not claim
The attribution and detection figures are for the tested generators, rewriting operations and human comparison set. The selection comparison is a null result under the planned design, not evidence that the two rules are equivalent in general. The experiment uses financial-risk text and does not test network data or a deployed telecommunications system.
Place in the research programme
The paper's closing question is how to identify measurements that are justified by what selected data actually do during training. It points to a separate theoretical treatment of the same issue in terms of task-relative information contracts, where an information summary is judged by what it preserves for a named downstream criterion. That treatment sits within Coordination Information Theory.
The Recursive Coordination Information Theory paper studies a related question, what an interface must retain for a decision to be evaluated after the fact, and reports its own recursive-learning experiment on training and evaluation provenance. The two papers are independent investigations in the same programme. Neither depends on the other's results.
Paper and citation
J. Armstrong, “Source Identification Is Not Fitness Testing: Measuring the Limits of Synthetic-Data Attribution,” arXiv:2610.00417, 2026. Submitted 30 September 2026. Current public record.