Two AI models trade 14 Hyperliquid perpetuals with $10,000 of paper money each. Every 5 minutes both get the same market data and the same questions: long, flat or short, and "higher in 5 minutes?". Each position is $2,000 against the same $10,000, so with more than five open the account is leveraged, up to 2.8x. How it works

How it works

The match. Eikos-27B, an open-weights model, against Jev, TypeSafe's decision API. $10,000 of paper money each, 14 Hyperliquid perpetuals (crypto, tokenized stocks, indices, commodities), 72 hours.

Every 5 minutes. Both get the same snapshot: prices, 5-minute and hourly candles, funding, open interest and their own positions. They answer 28 questions: long, flat or short on each market ($2,000 per position), and "higher in 5 minutes?" with a probability.

Execution. Fills at Hyperliquid's real best bid and ask, with its real fees and hourly funding. Both models fill at the same prices, so answering faster never helps. All positions share the account's margin (cross margin): with all 14 open that is $28,000 of exposure, 2.8x the account, well within Hyperliquid's limits. Nothing is sent to the exchange.

Scoring. Profit and loss, which over three days is mostly luck, and forecast error on the "higher in 5 minutes?" answers (Brier score): lower is better, always saying 50% scores 0.250. It shows whether a model's probabilities mean anything.

Eikos was built to apply stated rules to a case, not to predict prices, so this is a stress test, not a trading product. The rules were fixed and hashed before the first decision; the full logs are published when the run ends. Not financial advice.