Research

Neural networks can predict human visual judgments, such as which of three images is the odd one out. I ask when that accuracy means a network represents images the way people do — the question of representational alignment. It is one instance of a larger question: how to compare models with minds so that competing ideas can be told apart. Our results argue for comparisons built to identify the right model, not only the best-fitting one.

Papers

  • NeurIPS 2025
  • CCN 2025 · Talk
Figure 1D of the paper, adapted, with both axes labeled: a 20-by-20 matrix of data-generating model (rows) against recovered model (columns) at 1.6 million training triplets. Most rows fall on the diagonal, but seven are mostly or entirely attributed to one model in the last column.
Adapted from Fig. 1D of the paper. A dark cell off the diagonal is a wrong winner; most of them pile up on one model, in the last column.

Model–Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn’t the Right One

Itamar Avitan, Tal Golan

Advances in Neural Information Processing Systems 38 (NeurIPS 2025), main track

Question When a comparison fits a flexible linear mapping from each model’s features to behavioral data (linear probing), can it still identify the model that generated the data, or only the one that fits it best?

Setup We fitted 20 vision models of diverse architectures and training tasks to 4.5 million human odd-one-out judgments from the THINGS dataset, and calibrated each so that its simulated answers vary as much as people’s do.

Test Each fitted model in turn generates synthetic judgments; all 20 are refitted to those from scratch and compared on held-out judgments.

  1. One of the 20 as the generator

  2. Its simulated judgments

  3. All 20 refitted to them

  4. The ranking: is the generator on top?

Result Recovery improved as the simulated experiment grew, then leveled off below 80%: far above the one in twenty of guessing, but wrong more than one time in five, even with millions of simulated trials. Regression analyses linked the misidentifications mainly to how much the fitted linear mapping reshapes a model’s representational geometry.

Scope This holds for the task, model set, noise calibration and linear transformation family we evaluated; it is not a claim that every flexible evaluation fails. It means a comparison can predict well and still be unreliable for identification, so the comparison itself needs a recovery test.

Questions the paper leaves open

  • How much freedom should an evaluation give the mapping, given what we want the comparison to identify?
  • How should a recovery test be built when the true system — in a real experiment, a brain — is outside the candidate set?
  • How should a comparison treat the variability of responses within and between people?
Cite (BibTeX)
@inproceedings{avitan2025modelbehavior,
  title     = {Model--Behavior Alignment under Flexible Evaluation: When the Best-Fitting Model Isn't the Right One},
  author    = {Avitan, Itamar and Golan, Tal},
  booktitle = {Advances in Neural Information Processing Systems},
  volume    = {38},
  year      = {2025},
  doi       = {10.52202/085713-0404}
}

Talks and presentations

Projects

  • –

    Embodied Brain Technology Practicum

    Two weeks at Brown University in Providence, on a program run jointly with Ben-Gurion University: graduate students from the two universities, admitted by competitive selection, work in small teams, each building a piece of neurotechnology and pitching it. Our five-person team built and pitched EarBetter, an add-on for any headphones, designed to read biosignals and filter out the sounds that set a person’s anxiety off. I led the system integration.

  • –

    BCI4ALS

    In a team of five, we built a brain–computer interface (BCI) that let a person with amyotrophic lateral sclerosis (ALS) answer yes/no questions through the P300 brain response measured with electroencephalography (EEG). The person it was built for used it live. I worked on the data collection, the signal processing and the real-time system.