Research

ARCTIC

The Abstraction and Reasoning Corpus through Transfer & Induction Core — a visual reasoning benchmark for out-of-distribution generalization.

ARCTIC

In development

Rethinking OOD generalization in AI benchmarking through a transfer & induction core.

Current frontier benchmarks act as optimization drivers that push large models toward general-purpose behavioral imitation via expensive data annotation — lacking formal guarantees against test-set leakage. ARCTIC reframes evaluation.

Instead of testing unconstrained systems directly, it evaluates standardized small language models (SLMs) operating under a fixed Transfer & Induction Core. Transfer-learning mechanisms emulate abstract general capabilities; the constrained SLM is the actual evaluation target.

This closes the leakage loophole: AI companies cannot access private test inputs at the API level.

ARCTIC-0

A demonstration benchmark of 85 private ARC-AGI-style reasoning tasks, with an ambiguity-tagging mechanism. Tasks span pattern transformation, structural reasoning, object manipulation, and sparse completion — across easy / medium / hard / expert slices.

The trust model

ARCTIC analyzes both model-owner and benchmark-owner attack vectors — including ONNX transfer and the deliberate absence of model-owner authentication.

85
Reasoning tasks
7.05%
TRM baseline
12.94%
ARChitects
4
Difficulty slices

Explore

Public leaderboard, task catalog, and the paper live on the benchmark site.