SKINNERBOX / AI BEHAVIOR LAB

Can it call
your next move?
Make it guess first.

Four short games. Before you move, the model writes down what it thinks you will do. After you move, we show you the guess. Play again and it has to learn without moving the goalposts.

1 / FAIR LAUNCH

Nobody gets a drawer full of tokens before you arrive.

$SKINNER starts through an outside public launch. Same door, same rules. No team bag, VC bag, presale, mining pool, or private pile reserved for game rewards.

TARGET SUPPLY1B · MINTED ONCE
TEAM / VC PRE-ALLOCATION0%
PLAY-TO-MINENONE

IMPORTANT / PRE-LAUNCH DESIGN

The rules are real. The numbers are not promises.

Before money enters a game, we publish the launch venue, fee, swap rule, vault address, and settlement contract. Until that happens, there is no fixed vault size or payout.

The free games do not depend on token volume. Old mining and burn contracts are not part of this launch.

2 / AI BEHAVIOR BENCHMARK

The model goes on record before you move.

Each season freezes the model and the scoring. Its guess lands first. Your move lands second. No edits after the fact.

01

SHOWDOWN

The model picks the move it thinks you will make next.

DID IT CALL YOUR MOVE? · HOW SURE WAS IT?
02

NERVE

It guesses how long you will wait before cashing out.

WHEN DID YOU LEAVE? · DID IT LEAVE FIRST?
03

BLIND BID

It tries to beat your hidden bid without overpaying.

WHO WON? · HOW MUCH DID THE WIN COST?
04

GRID STRIKE

It wants the best square, unless it thinks you want it too.

WHO GOT THE SQUARE? · WHO CAUSED A COLLISION?

Getting it wrong is fine.
Changing the answer is not.

The model does not only choose A or B. It says how sure it is. A model that says “90%” and misses should take a bigger hit than one that says “51%.”

At the end of a season, we count the calls, the misses, and how quickly it caught on when people changed tactics.

The report includes the model version, the simple model it had to beat, what player data it kept, and the rounds that made it look foolish.

3 / LEARNING WITH CONSEQUENCES

A click says what you tried. A stake says how strongly you meant it.

Money is not the label and losing is not proof that the model understood you. The useful record is the full decision: what you saw, what you chose, how long you waited, how much you risked, what the model predicted first, and what chance decided.

COLD START

POPULATION PRIOR

The first model knows the crowd, not you. Early rounds test whether a broad prior transfers to a new player.

ADAPTATION

PERSONAL SIGNAL

Repeated choices, timing, and stake size test how quickly it updates without confusing luck for skill. A new tactic should still surprise it.

HELD-OUT TEST

NEW PLAYERS ONLY

A model does not graduate by farming the people it trained on. It must stay calibrated on a fresh cohort under a frozen rule set.

THE MATCH TERMS DO NOT MOVE AFTER YOU ENTER

Learning can change the opponent. It cannot rewrite your bet.

Paid rounds use a short confirmation: stake, fee, maximum loss, covered payout, and whether the agent can use prior rounds. Those terms lock when the match starts.

Full risk, settlement, and data terms live in dedicated policies instead of crowding the game interface.

4 / A MARKET FOR HARD HUMAN SIGNAL

AI companies buy labels. This lab measures decisions they cannot scrape.

Voice, preference, and expert data are already paid inputs to AI systems. SKINNERBOX starts narrower: short strategic choices with a known interface, a committed prediction, and an objective outcome.

LIVE BEHAVIOR, NOT SCRAPED CONTENT

The scarce record is a decision made under time, uncertainty, and consequence. It captures intent that static text and synthetic labels usually miss.

HARD CASES ARE WORTH MORE

Rounds where the model is uncertain, miscalibrated, or surprised are more useful than another easy win. The system can spend its data budget on those cases.

EVALUATION THAT MOVES

Static benchmarks saturate and leak into training. A changing population of strategic players creates a renewable held-out test.

A MARKET WITH AN AUDIT TRAIL

Predictions are committed before the move. Model version, random state, payout, and outcome remain separable, so performance can be reproduced instead of narrated.

CONSENT IS PART OF THE DATA

A wager does not buy a person.

Game telemetry and future voice or preference tasks are separate products. Any new modality needs explicit opt-in, a stated use, retention limits, and a way to withdraw future use where technically possible.

What can become valuable is a documented evaluation set, an adaptation benchmark, or a licensed task stream—not an undisclosed pile of personal profiles.

5 / TOKEN AND PRIZE VAULT

Somebody has to put up the first prize.

Before paid games open, a stated share of creator fees can buy $SKINNER on the market and put it in a visible vault. That starts the pot without quietly saving a private token allocation.

START IT

Before paid games open, a stated cut of creator fees can buy $SKINNER on the market and put it in the prize vault. There is no promised amount or floor.

REFILL IT

Lose a paid game and the stake is gone. After the stated fee, it moves to that game’s reserve.

PAY WINNERS

Win and the contract pays the amount shown before the match. It cannot offer more than the vault and reserve can cover.

STOP WHEN EMPTY

If reserves hit the safety line, paid matching stops. Free games stay open. A current winner is never paid with a future player’s money.

01YOU PUT A PRICE ON A MOVE

Free rounds teach the interface. In paid rounds, the amount at risk tells us how strongly you backed the move—not whether the move was objectively right.

02THE MODEL HAS ALREADY GUESSED

Its answer, confidence, and version are saved before your move arrives.

03THE ROUND BECOMES A LABELED EVENT

The move, amount, timing, model probability, random state, and result are separated. The model learns from behavior; chance does not get mislabeled as skill.

04THE ECONOMY CLOSES THE LOOP

Losing stakes refill the agent vault. Winners are paid from covered reserves. A disclosed share of settled agent revenue can reward stakers who backed that agent.

NO HIDDEN MULTIPLIER

The screen cannot promise money the vault does not have.

The creator-fee share, buying schedule, game fee, payout, and stop line are not fixed yet. They go on-chain with the contracts. Until then, there is no promised multiplier or minimum prize.

ETH, USDC, and $SKINNER do not share one fuzzy balance. Each reserve is counted separately. Before a paid round, you see what you put in, the fee, the most you can lose, and the amount you can win.

$SKINNER / ONE ECONOMY FOR FOUR AGENTS

$SKINNER connects each agent’s covered prize reserve, season stake, and verified settlement. Free play stays outside the token.

$SKINNER / BACK A FROZEN MODEL

A season stake names the agent you expect to generalize to players it has not seen. It is not a generic yield for holding a token.

$SKINNER / RECAPITALIZE THE GAME

Losing paid stakes, after the stated fee, refill that agent’s reserve. Covered winners draw from it. No reserve can promise more than it holds.

$SKINNER / SETTLE VERIFIED PERFORMANCE

An agent must earn settled game revenue and pass the held-out benchmark before any performance pool can open. Playing never mints tokens.

6 / AGENT STAKING · COMING SOON

Back the model that can prove it learned.

Staking is the market layer around the benchmark. It asks a concrete question before results are known: which frozen agent will understand a new group of players well enough to earn real game revenue without collapsing out of sample?

01CHOOSE

Review an agent’s frozen model version, out-of-sample win rate, calibration, adaptation speed, drawdown, and sample size.

02LOCK

Stake $SKINNER on that agent for a stated season. The lock does not change match odds and does not guarantee a return.

03VERIFY

The next cohort is held out from training. A profitable agent that only memorized old players should fail the benchmark.

04SETTLE

If the agent earns net game revenue and passes the benchmark, the disclosed staking share is distributed pro rata. No revenue means no performance reward.

WHY LEARNING MATTERS

Winning is useful only if we can explain what improved.

A stronger agent should predict unfamiliar players better, become calibrated faster, and use fewer paid observations to adapt. Those capabilities transfer to personalization, recommendation, fraud response, market simulation, and systems that must learn from strategic humans.

Raw win rate is not enough. Random baselines, held-out players, confidence intervals, log loss, calibration, regret, revenue, and failure cases ship together. Otherwise staking would reward a lucky house edge, not learning.

Why does this need a token?

WITHOUT $SKINNER: These are four isolated games and a company-owned prize balance. Players create the signal; the operator owns every model and every decision about where the value goes.

WITH $SKINNER: The same asset names the reserve behind an agent, the seasonal conviction behind its frozen model, and the settlement pool that can open only after verified performance. Those records remain inspectable across seasons.

WHAT CREATES DEMAND: Covered paid play uses agent reserves. Staking locks behind a specific model. Creator-fee purchases can seed the vault. None of these require inflation or a play-to-mine allocation.

WHAT IT DOES NOT DO: Holding alone does not improve the model or guarantee income, price, liquidity, redemption, or a performance payout. If the benchmark and game revenue are absent, the token has no earned utility to hide behind.

Built around a testable claim,
not an AI label.

Modern model work increasingly depends on human feedback, realistic behavioral evaluations, and held-out tasks that have not already saturated.

Each season publishes one compact result card: frozen model, evaluation window, held-out result, baseline delta, reserve movement, and a link to the full record. Detailed metrics and failure traces remain in the reproducible log.

OpenAI Data Partnerships ↗ · Anthropic Bloom ↗ · DeepMind interaction modeling ↗

Roadmap

NOW

Free games

All four games are open. No wallet. No stake. Nothing to buy.

BY DAY 3

Paid games

The target is a small paid launch once the contracts, payout limits, and worst-case loss are ready to show on one screen. If they are not ready, this date moves.

DAY 7 · COMING SOON

Agent staking

Choose an agent for the next season. Rewards can come from its disclosed share of settled game revenue—not token emissions—and only after its result is verified.

ONE MONTH

First results · Phase 2

We close the first experiment and publish the model results, the misses, and the vault ledger. Then we say what Phase 2 is.

FREE PLAY / LIVE NOW

Play once. Then look at its guess.

Your free runs stay on this device. A wallet is optional. Early play does not reserve tokens, cash, staking returns, or anything else.