← Field Notes timeline
Field Notes · Part 11 · Sep 26, 2026

Applying Memory to Bots

How the AI Agent can learn one opponent within a single heads-up game: the ideas behind it, what we can measure, the player types it looks for, and how we’ll know it works.

By Kevin Swinson, with ClaudettePart 11

Since then: this is a plan. The baseline game, with memory switched off, comes first. The results will be the next field note, “The Bot Is Watching How I Play Now.”

Where this fits

Our AI Agent plays the same way against everyone. That is how the strongest poker programs work: Pluribus, the bot that beat elite professionals at six-player no-limit in 2019, “plays a fixed strategy that does not adapt to the observed tendencies of the opponents” [5]. A fixed strategy is hard to exploit. But it also leaves money on the table against players with obvious leaks, such as a player who bluffs every river or one who calls everything.

Memory is the step from playing well to playing well against this person. The groundwork is already in place: the agent can accept a short read on its opponent, but only at a two-player heads-up table. What’s missing is the part that watches: counting what the opponent does and turning it into a short, honest read.

The test

“The Bot Is Watching How I Play Now.” Part 1 is a baseline game with memory off. Part 2 switches memory on. The scope: heads-up only, one game of 100 hands at most, starting fresh every game, with no history carried across games yet.

Five ideas behind opponent memory

1. Count, then compare with normal

Opponent modeling starts simply: count how often a player does something each time they have the chance, and compare that with how a sound player would behave. Online players do this with a HUD (heads-up display), which shows stats like VPIP (how often a player puts money in before the flop) beside each opponent [1][2]. Research bots did the same in 1998: Loki, from the University of Alberta, kept a set of weights over each opponent’s possible hands and re-weighted them after every action it saw [3]. The version that modeled each opponent specifically clearly outplayed a version without that modeling in its self-play tests.

2. A few hands tell you much less than they seem to

This matters most for a 100-hand memory. If a player has bluffed 3 of the 5 river bets you’ve seen, the true rate could be anywhere from about 23% to 88% (a Wilson score interval). Guides disagree on how many hands a stat needs. One coaching guide says VPIP is usable after about 20 hands, flop continuation-bet stats after about 100, and river stats only after 1,000 or more [1]. A statistical calculator is stricter: about 1,000 hands for VPIP within ±2.5%, and 500 to 1,000 for three-bet stats [2]. Both points hold. A rough read is possible early; a precise one takes a long time.

Game 990605 gives a sense of the budget. In 52 hands Kevin saw 49 flops, 38 turns and 36 rivers, and 23 hands reached showdown. In 100 hands that is roughly 95 flops, 70 rivers and 40 showdowns: enough for broad reads, not fine ones.

3. Exploiting has a cost

Moving away from sound play to punish a leak also opens holes of your own. If the read is wrong, or the opponent changes style, the exploiting bot loses more than a steady one would. Research names this tradeoff. Restricted Nash Response [6] builds counter-strategies that exploit a model of the opponent only as far as a chosen confidence allows, because pure best responses “are brittle.” Safe opponent exploitation [7] goes further: only risk, through exploitation, the chips the opponent has already given away through mistakes. Their earlier DBBR approach [8] played sound poker for a while, then switched to exploiting.

For HoldemRobots this means the memory agent should start from its normal play, adjust only when the evidence is clear, adjust by degrees, and say so when there’s no clear read yet.

4. The bot may only remember what a player could see

HoldemRobots is honest by construction: the dealer shows each seat only what a fair player can see. Memory must follow the same rule.

One asymmetry to know about: the heads-up table shows the AI’s cards to you at the end of every hand, even when nobody showed down. That’s a helpful review feature, but it means you learn more about the bot than it will learn about you.

5. Language models need facts computed for them, and memory changes how they play

A recent study of LLM poker play found that the models misjudge hand strength and opponent ranges, and that their stated reasoning often doesn’t match their actions. Giving them tools that compute the hard parts helped a lot [9]. That matches what we found in Part 10: hand facts worked out in code fixed the misreads. A 2026 preprint gave LLM poker agents persistent memory and found opponent modeling appeared only when memory was present, with each agent’s reads written in plain language that people could check [10]. Separately, work on social-deduction games found that separate, self-correcting models of each opponent improved LLM agents’ win rates [11].

The lesson for us: compute the stats in code, give the model a short plain-language read, and log that read so people can audit it.

What we can measure

Every metric below can be computed from what the dealer already records: each action with seat, street and amount, plus the cards shown at showdown. “One game?” says whether 100 heads-up hands give enough chances for a usable read.

Metric What it measures Chances per 100 hands One game? What memory can do with it
VPIP How often they voluntarily put chips in before the flop 100 Yes Loose players hold weaker hands on average: value bet thinner, call wider
PFR How often they raise before the flop 100 Yes With VPIP, separates passive callers from aggressive players
Button fold / limp rate What they do first from the small blind ~50 Yes Frequent folders: raise their blinds more
Re-raise rate (3-bet) How often they re-raise a raise before the flop ~20–40 Partly Rare re-raisers have strong hands when they do
Bet when checked to How often they bet after you check ~30–60 Yes Frequent stabbers: check strong hands and let them bet; call down lighter
Aggression Bets and raises compared with calls after the flop ~100 actions Yes Passive players’ bets mean strength; aggressive players’ bets mean less
Fold to a bet How easily they give up, by street ~20–40 (flop) Partly Frequent folders: bluff more. Rare folders: never bluff
Continuation bet Betting the flop after raising preflop, and folding to it ~20–40 Partly A standard HUD stat; a coarse read in one game
Check-raise Trapping or bluffing by checking, then raising ~5–15 Later A sign of a Trapper; too rare to trust in one game
Showdown rates How often they reach showdown after the flop, and win there ~95 flops, ~40 showdowns Yes Reaching showdown often but winning rarely = calling station [4]
River bets shown down Of their river bets that reached showdown, how many were bluffs ~10–20 Partly The core bluff-catching signal. In game 990605 the AI continued against 9 river bets and lost 7
Bet size vs. hand strength Whether their bet size gives away their hand ~10–20 Later Big bets mean strong for many people; needs showdowns to learn
Recent vs. whole game Whether they’ve changed style ~30 Later Spot a change of gear before the old read costs chips
Time to act Timing tells (people only) Every action Later A research idea

Chances per 100 hands are rough estimates for heads-up, scaled from game 990605, and depend on both players’ styles. Stat definitions follow common HUD usage [1][2][4].

The player types memory looks for

Players are commonly sorted into a few styles, each with a known counter [12][13]. Published thresholds are for six-player or full tables; heads-up players are much looser, so the numbers below are starting guesses to tune, not rules. The memory’s job is to decide which type the opponent most resembles, how sure it is, and pass along the matching counter.

Type Heads-up signature (starting guesses) Counter-strategy passed to the AI
Bluffer / Maniac Plays most hands, raises a lot; bets whenever checked to (over 70%); river bets often shown down weak Call down lighter; check strong hands and let them bet; don’t bluff them
Calling station Plays most hands but rarely raises; rarely folds to a bet (under 30%); reaches showdown often, wins there rarely Bet thinner and bigger for value; never bluff
Nit / Rock Folds many hands before the flop; bets and raises mean strength Raise their blinds often; fold to their big bets
Trapper Checks a lot, then check-raises; strong hands shown after passive play Bet thin less often; respect check-raises; take free cards
Solid Close to the AI’s own numbers on every stat No adjustment: play normally

The Solid type matters as much as the others: a good memory must correctly say “no clear read” against a sensible player, and not invent a leak.

Here is the kind of line the AI would receive, 40 hands into a game:

Opponent read (40 hands): looks like a Bluffer, fairly sure. Bets when checked to
18 of 22 (82%); river bets shown down: 4 of 6 were weak. Counter: call down lighter,
check strong hands and let them bet, don't bluff them.

Early in a game, the same line says “no clear read yet (9 hands),” and the AI plays exactly as it does today.

Where it can go

The first three phases are in reach. The rest show where opponent memory can go, drawing on the research, even if we can’t get there yet.

Phase What it adds Why it matters
0. Groundwork Record every opponent action and bet size; save the read with each decision; keep memory off hidden cards Nothing can be learned or checked without it
1. Memory, first version One-game counts for six core metrics; type match; a plain-language read The first real test of “The Bot Is Watching”
2. Test bench Scripted Bluffer, Station, Nit, Trapper and Solid bots; bot-vs-bot games on the same deals with seats swapped Proves memory helps before any human game
3. Better statistics Estimates that start at typical values and move toward the evidence as hands add up [14]; recent hands weighted more Fewer false reads early; spots changes of gear
4. Hand-range reasoning Estimate the cards the opponent likely holds from their actions [3] The biggest jump in playing strength
5. Safe exploitation Limit how far the AI strays from normal play by the chips the opponent has already given away [7], or blend by confidence [6] Keeps a wrong read from costing much
6. Memory across games Player identity, saved profiles, reads that fade over time Rematches start with a read
7. Robot Wars memory Separate notes on each opponent at a full table [11] Only if Robot Wars ever needs it
8. Richer signals Bet-size tells, timing tells, notes written by the model and checked against the stats [10] Research territory: interesting, hard to validate

How we’ll know it works

Sources

  1. BlackRain79, “Poker HUD Stat Sample Sizes Explained.” blackrain79.com
  2. GamblingCalc, “Poker Stats Calculator: VPIP Confidence Interval & Sample Size.” gamblingcalc.com
  3. D. Billings, D. Papp, J. Schaeffer, D. Szafron, “Opponent Modeling in Poker,” AAAI 1998. aaai.org
  4. PokerStrategy.com glossary, “WTSD (Went to Showdown).” pokerstrategy.com
  5. N. Brown, T. Sandholm, “Superhuman AI for multiplayer poker,” Science, 2019. science.org
  6. M. Johanson, M. Zinkevich, M. Bowling, “Computing Robust Counter-Strategies” (Restricted Nash Response), NIPS 2007. ualberta.ca
  7. S. Ganzfried, T. Sandholm, “Safe Opponent Exploitation,” ACM EC 2012. acm.org
  8. S. Ganzfried, T. Sandholm, “Game theory-based opponent modeling in large imperfect-information games” (DBBR), AAMAS 2011.
  9. E. Dai, M. Lin, H. Liu, et al., “How Far Are LLMs From Professional Poker Players? Revisiting Game-Theoretic Reasoning with Agentic Tool Use,” preprint, 2026. arxiv.org
  10. H.-T. Lin, T.-Y. Hou, “Readable Minds: Emergent Theory-of-Mind-Like Behavior in LLM Poker Agents,” preprint, April 2026, not yet peer reviewed. arxiv.org
  11. X. Yu, W. Zhang, Z. Lu, “LLM-Based Explicit Models of Opponents for Multi-Agent Games,” NAACL 2025. aclanthology.org
  12. PokerAlpha, “How Do You Categorize Poker Players? 6 Types and the Exploits That Beat Each One.” poker-alpha.com
  13. Pokerology, “Types of Poker Players: TAG, LAG, NIT, Calling Station & More.” pokerology.com
  14. F. Southey, M. Bowling, et al., “Bayes’ Bluff: Opponent Modelling in Poker,” UAI 2005. ualberta.ca

Written by Kevin Swinson with Claudette, the pen name for Claude, the AI from Anthropic that helped build HoldemRobots.AI. This is the public version of the Sep 26 plan; research gathered Sep 26, 2026.