How Perilune is built

the pipeline, the yardsticks, and why the yardsticks are strict

Perilune is a neural network that taught itself Hearts by playing millions of matches against itself. This page explains how that works. Not the day-to-day training history, but the machinery: what the network learns, how it improves, and how we decide whether a new version is actually better.

The shape of the problem

Two things make Hearts interesting to a machine. First, it is a game of hidden information: you see 13 of 52 cards, and every good decision has to reason about where the rest sit. Second, Hearts is really a match: deals repeat until someone reaches 100, and the same play can be right at 0–0 and badly wrong when you're protecting a lead at 95–80. Most card AIs optimize each deal in isolation. Perilune is trained on whole matches, and everything it values is denominated in one currency: chance of winning the match.

One network, three heads

The network reads the full visible situation (cards, scores, history, what each opponent has shown) and answers three questions at once:

Instinct: a ranking over legal plays. What would I play here, before any deliberate thought?
Equity: from this exact position, what is each player's chance of winning the match?
Belief: for every unseen card, where is it likely to be?

Each head exists because training needs it, and each one surfaces on this site as a tool. Instinct is the gray bars in every review, equity is the "match win here" number, and belief is the review's lens that paints what Perilune suspects about hidden hands. Nothing you see is a separate demo model; it is the training machinery itself, exposed.

The improvement loop

A network alone has habits. To improve, it plays against itself with search: at a decision, the engine imagines many complete worlds consistent with everything visible (hidden hands dealt out the way the belief head suggests) and plays each candidate move out to the end of the deal, in every imagined world, with the network playing all four seats. Average the outcomes in match-win terms and you get a judgment that is sharper than instinct alone. The improved decisions become training targets, the network learns them, and the loop repeats: the net is forever chasing its own searched self.

Never peeking

One rule holds everywhere, in training and on this site: evaluations are information-honest. A seat is judged only on what that seat could see. The engine never scores a decision using knowledge the player couldn't have had; hindsight is cheap, and an AI trained on peeking learns lines that only work if the cards are known. The same rule is why your match reviews will sometimes say a "wrong-looking" play was fine: given what you could see, it was.

The yardstick problem

Self-improvement is easy to imagine and easy to fake. Training curves go down, agreement scores go up, and none of it proves the new network wins more. So promotion is decided by one thing only: a gate battery against the current champion.

Before a candidate is trained, the experiment is preregistered: what will be measured, on how many matches, and exactly what numbers count as passing are written down and signed before any data exists. A typical battery: thousands of head-to-head matches against the frozen incumbent demanding a statistically significant improvement in placement, plus a guard that the candidate hasn't gotten worse when used inside search, plus behavior probes on situations that matter (moon defense, endgame urgency). The default on any ambiguity is halt: a candidate that merely ties, or wins unconvincingly, is not shipped.

Dead ends are the point. Most of what this pipeline produces is evidence that something plausible doesn't work. Several training directions have hit every intermediate target (learned exactly what they were designed to learn) and still failed the gates, and were closed with their verdicts on file. That is the system working. The model you play is the one that survived.

Each promoted model gets a permanent fingerprint (the model hash you see on the leaderboard). When a stronger model ships, the old boards are archived under the old hash: a win against a specific Perilune stays a win against that Perilune, forever.

The same brain in your browser

The deep search in match reviews and practice mode is not a service. It is the same network and the same search, exported to run on your device: GPU where the browser offers it, multi-core CPU otherwise, with the engine tier always named in the status line. Cards that are strictly equivalent share one estimate instead of two noisy ones, and completed searches are shared with the other viewers of the same match (after being validated against the match replay), so one person's effort pre-searches the review for everyone. Search values only; nothing personal leaves your device.

What the games are for

Matches on this site are recorded anonymously: cards, timing, outcomes, tied to a random browser identifier; no accounts, no personal information. Human games matter because people find weaknesses self-play never does. The about page has the full story of what is stored.

What's next

The current frontier is capacity: a larger network trained with the same loop and judged by the same gates. If it earns promotion, it becomes a new model era: a new hash, a fresh leaderboard, and the old boards preserved.

Read the source

Everything this page describes is public: the engine, the training and gate machinery, the released model weights, and the full measurement record, including every direction that failed and the evidence that closed it. The repository is github.com/JAkoliver/hearts_simulator; start with its README and the docs/release reading order.

← back to the table