Can Circuit Alignment Predict OOD Generalization?
TL;DR: We introduce Circuit Alignment Score, a weight-only metric that predicts a model’s OOD generalization without target-domain data or labels, with strong theoretical guarantees and substantially higher correlation with OOD accuracy than existing representation-based metrics.
Circuit Evolution prediction under distribution shift with CAS: From weights alone, we extract class-specific circuits and compare them via graph kernels to score and rank learners based on OOD accuracy, without target data or labels.
Abstract
Can out-of-distribution (OOD) generalization be predicted from a trained model’s weights alone, without any target-domain data? Existing representational similarity metrics (CKA, SVCCA, RSA) compare activations rather than forecast generalization. We show they are provably insensitive to structural rerouting in the computational graph, the very change distribution shift induces. We close this gap with the Circuit Alignment Score (CAS), which compares class-specific circuits across domains via graph kernels, decomposed into same-class coherence and cross-class confusion. Casting CAS as a Lebesgue integral over the domain distribution, we prove its Monte Carlo estimate recovers the ground-truth ranking of learners by OOD accuracy, with pairwise inversion error vanishing at rate O(1/M), where M is the number of sampled domains. Across 48 learners on PACS, CAS attains 0.88 rank correlation with OOD accuracy, versus 0.58 (CKA), 0.23 (SVCCA), and 0.14 (RSA), with similar trends on other benchmarks and even against data-dependent methods, making it the first provably consistent predictor of distributional robustness requiring neither target-domain data nor labels.
Main Findings
CAS vs OOD accuracy shows a monotonous trend (ρS=0.93), whereas CKA (ρS=0.58), SVCCA (ρS=0.23), and RSA (ρS=0.14) show a diffuse scatter, confirming that representational similarity fails to capture the circuit-level signal (Target class: cartoon of the PACS dataset).
Consistency of CAS Rankings
Consistency of CAS rankings: Effect of the number of source domains M on ranking stability. Top: Pairwise inversion probability Pinv(M) vs. M (mean ± std over 10 trials) with O(1/M) Chebyshev bound (dashed lines). Bottom: Spearman ρS between M-domain CAS ranking and true OOD accuracy. TK consistently achieves the highest ρS and all methods follow the O(1/M) trend.
BibTeX Citation
@misc{banerjee2026circuitalignmentpredictood,
title={Can Circuit Alignment Predict OOD Generalization?},
author={Ayan Banerjee and Abhra Chaudhuri and Josep Llados and Umapada Pal and Anjan Dutta},
year={2026},
eprint={2609.31996},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2609.31996}
}