about
Training Leaves Traces: Centered Residual Signatures for LM Lineage Verification (arxiv.org)
2 points by singh96aman 46 days ago | hide | past | pdf | discuss on HN

In plain words: It strips away the weight pattern shared by all trained models and compares what's left, so weights alone can show if two models are related. The check perfectly separated descendants from unrelated ones, even after disguise, and runs 76x faster than the closest rival.

Abstract · Training Leaves Traces: Centered Residual Signatures for Language Model Lineage Verification

Open-weight language models are fine-tuned, quantized, pruned, and merged, yet their provenance is often undocumented. We study data-free white-box lineage verification: can weights alone reveal whether two compatible model checkpoints share ancestry? Residual training produces a shared identity-aligned component in branch products, so this structure alone cannot establish ancestry. We remove it and compare checkpoint-specific structure across residual blocks, yielding a symmetric lineage score calibrated against independent checkpoints. On residual-MLP and GPT-2 benchmarks, the score separates fine-tuned, LoRA-merged, pruned, and quantized descendants from independent and distilled models (AUROC=1.0), distinguishing weight ancestry from behavioral similarity. Under function-preserving checkpoint laundering experiments, weight-space baselines lose margin or fail; our score remains unchanged and runs 76x faster than the nearest robust baseline on GPT-2. The projection-pairing signal appears across six language-model families and beyond, and a case study correctly identifies 3 related and 7 unrelated LLaMA-2 public checkpoints. Collectively, these results establish a passive, data-free provenance signal for compatible open-weight language-model checkpoints

Aman Singh Thakur, Rayan Khoury
arXiv:2608.14929 · cs.CL, cs.LG · submitted Aug 14, 2026 · updated Sep 22, 2026
abstract · pdf · html · Preprint

add comment on HN