← All articles

Kaiming He's team shows an AI pretrained on cat photos solves ARC almost as well as LLMs

Glowing glass panels divided into colored cells, some cells highlighted

Kaiming He's team — the author of ResNet and Masked Autoencoder, now a professor at MIT — has published a paper for ECCV 2026 in which a vision model pretrained on ordinary ImageNet photos (cats, dogs, flowers) solves tasks from the abstract ARC benchmark: 63.4% pass@2 for a single model and 70.2% for an ensemble. There is no text and no language in this model — only "vision" transferred from real photographs to abstract colored grids. It is the first time a purely visual approach has come this close to specialized LLM systems, and it did so with a quarter of their parameters.

What ARC is and why it is hard

ARC (Abstraction and Reasoning Corpus) was created by François Chollet, the creator of Keras, in 2019. The format is uniform: a grid of up to 30×30 cells, up to 10 colors, and 2 to 5 input-output pairs. From the examples you must infer a hidden transformation rule and apply it to a new grid. It sounds like a children's pattern puzzle, but for machines it is one of the harshest tests of abstract thinking: the rule is new every time, there is no template to memorize, and there are only two or three examples.

How ARC is solved today

Language models dominate: the grid is translated into text or a symbol sequence and handed to an LLM for reasoning. The visual route lagged behind for a long time, but it is exactly where the most interesting movement is coming from. The same MIT group, in their earlier VARC work, redefined ARC as a conditional image-to-image translation task and pushed a 19M-parameter model to 54%. Refinements followed: LoopViT with recurrent reasoning — 18M parameters and 65.8%; Loop-OWM with video pretraining — 10.6M and 68.5%. Orders of magnitude fewer parameters than LLM systems, and the results are already close.

What visual models lacked was pretraining. Language models scale because they are pretrained on gigantic text corpora, while visual ARC solutions started from random initialization and never got that bonus.

NAT-ARC: pretraining on cat photos

NAT-ARC inserts an MAE pretraining step on ImageNet — a dataset of roughly 1.3 million natural photographs — in front of the VARC visual pipeline. MAE (Masked Autoencoder) is He's own method from 2022: the model masks most of an image and learns to reconstruct the original from the remaining fragments, picking up shape, structure and spatial relations along the way.

A practical detail: the researchers took a ready-made public MAE checkpoint and spent nothing on pretraining. The checkpoint was built for large ImageNet images while ARC grids are only 64×64 pixels, so they kept only the backbone, discarded the patch embedding and positional encodings, and installed a flexible 2D RoPE instead.

Glowing glass panels divided into colored cells, some cells highlighted

Why cats help solve abstract tasks

The authors placed three models with different initializations — no pretraining, ImageNet MAE, and ARC-style MAE — in front of ARC tasks before any training and looked at their attention maps. The random model's attention is spread uniformly, with no focus. The ImageNet-pretrained model's attention already concentrates on meaningful regions: it knows how to separate an object from its background, although it has never seen a single ARC grid.

The researchers picked 15 tasks where pretraining gave a clear boost and annotated them by hand. Six turned out to be match-and-copy tasks — recognize an object, color or pattern and copy its structure to the output. Another six were connected-component reasoning — isolate, fill or recolor a connected region of the grid. "Separate the cat from the grass" becomes "separate a connected colored region from the background" — the visual and abstract worlds share the same underlying mechanics.

The numbers: 0.6B parameters versus 8B

On ARC-1, the best single NAT-ARC model at the huge scale — 0.6 billion parameters — scored 63.4±0.7% pass@2. An ensemble of three models with different pretraining strategies (majority voting, about 2B parameters in total) reached 70.2±0.6%. For comparison: the best specialized LLM system, The ARChitects, scores 71.6% — but on an 8-billion-parameter language model. The visual route closed in on it with a quarter of the parameters.

Pretraining fixed scaling

The second important result is about how models grow. Without pretraining the curve breaks: accuracy rises from base to large but falls from large to huge — a bigger model with no better results, because ARC has too little data and the extra capacity goes into overfitting. With ImageNet pretraining the curve turns healthy: all three scales improve together with size. That is unlocked scaling for the visual route.

The authors got an elegant proof from the other direction: an autoencoder maps an ordinary ImageNet photo into a discrete 60×60 grid with 10 colors — essentially "translating a cat into the language of ARC." NAT-ARC applies its learned transformations to that grid — a 180° rotation, a blue border — and decodes it back into pixels. The cat really did rotate, and the border really did appear. Abstract rules and visual representations live in one shared space.

Who is behind it

The first author is Xiaoman Delores Ding of MIT CSAIL, from He's group; she led the earlier VARC as well. Then Keya Hu, Katelyn Gan and Victor Yin — also MIT. Co-author and corresponding author is Kaiming He himself; through ResNet and MAE you can trace almost a decade of visual AI. Amusingly, in this new work he uses his own four-year-old method as a tool: an old weapon opened a new road.

What it means in practice

The main takeaway: size is not everything — pretraining decides. A compact model with good preliminary training catches up with a system four times its size. Image generation follows the same logic: the winner is not the largest model but the one best trained for the task and cheaper to run. According to our service's data, over the last 30 days (11,177 image generation attempts from September 2 to October 1, 2026) the most in-demand model was by no means the heaviest one — nano-banana-2 accounted for 2,560 completed generations, about 23% of all, at an average cost of roughly $0.05 per image (the generation_costs table for the same period). Reliability of compact models is fully production-grade: over the same month roughly every thirteenth attempt failed to produce a result — 829 failures out of 11,177.

An honest caveat: ARC-1 is a single benchmark, and universal visual reasoning is still far away. But the direction is clear: instead of endlessly inflating language models, you can train vision on the real world and transfer it to abstractions.

FAQ

What is ARC?

An abstract reasoning benchmark by Keras creator François Chollet: from 2–5 input-output examples on a colored grid, infer a hidden rule and apply it to a new grid. It is considered one of the hardest benchmarks for AI.

Why cats specifically?

The point is not cats but pretraining on natural images — ImageNet is full of them. A model that learned to separate objects from backgrounds on photos transfers that ability to abstract grids.

Can I try image generation?

Yes, in the NeuralSpace image generation section — it gathers models of different classes, from fast ones to maximum quality.