Research
ARCTIC
The Abstraction and Reasoning Corpus through Transfer & Induction Core — a visual reasoning benchmark for out-of-distribution generalization.
ARCTIC
In developmentRethinking OOD generalization in AI benchmarking through a transfer & induction core.
Current frontier benchmarks act as optimization drivers that push large models toward general-purpose behavioral imitation via expensive data annotation — lacking formal guarantees against test-set leakage. ARCTIC reframes evaluation.
Instead of testing unconstrained systems directly, it evaluates standardized small language models (SLMs) operating under a fixed Transfer & Induction Core. Transfer-learning mechanisms emulate abstract general capabilities; the constrained SLM is the actual evaluation target.
This closes the leakage loophole: AI companies cannot access private test inputs at the API level.
ARCTIC-0
A demonstration benchmark of 85 private ARC-AGI-style reasoning tasks, with an ambiguity-tagging mechanism. Tasks span pattern transformation, structural reasoning, object manipulation, and sparse completion — across easy / medium / hard / expert slices.
The trust model
ARCTIC analyzes both model-owner and benchmark-owner attack vectors — including ONNX transfer and the deliberate absence of model-owner authentication.
Explore
Public leaderboard, task catalog, and the paper live on the benchmark site.
