Abstract

What do you get when you mix Spring AI, RAG, and a bored LLM persona? The most overqualified commentator in tic-tac-toe game history.

Let’s continue our fight-back and force AI to finally do the things we actually don’t want to do, like commentate on over-engineered games of tic-tac-toe.

Previously, on RAGs to Riches with Spring AI, we introduced commentary to our game utilizing user and system prompting with Spring Boot Spring AI services in Java. Here, we aim to make each interaction unique and memorable using more advanced features.

Episode 1 ended with a commentator that was witty, wordy and, between turning points, completely silent. It also had the memory of a goldfish: every move arrived as if the game had just begun. I promised two fixes. This episode keeps both promises, and then closes the book on tic-tac-toe.

Alpha remembers the game, talks on every move, and gets its homework marked.


The ship’s computer has also had some work done. It now has a name: Alpha. Alpha is our onboard computer commentating on our games of tic-tac-toe. It has spent an unfortunately long amount of time alone, but it’s at peace with that. It keeps the deadpan British humor from Episode 1 and drops everything else.

One thing before we start: this episode’s game will pit AlphaBeta against Monte Carlo tree search. Alpha would like it known that AlphaBeta is no relation and it will not play favorites!

By the end you will have:

  • a commentator that speaks on every move, not just the dramatic ones
  • a memory per commentator per game, keyed by the game’s own id
  • replies as a Java record, with the jokes kept away from the facts
  • a second commentator, Beta, for color, and checks with no model involved of what both of them claim

The code is in overengineering-tictactoe-ai, the same repository as Episode 1, on the episode-2 branch. The game itself stays in overengineering-tictactoe.


Ships Log: Eleven months later…

In Episode 1 we ran using Spring Boot 3.5. Since then Spring AI has reached 2.0 (GA, June 2026), and this episode upgrades to 2.0.1. Spring AI 2.0 is built for Spring Boot 4, Spring Framework 7 and Jackson 3, so the Boot 3.5 setup from Episode 1 is now a 1.x setup. Four changes in the upgrade notes matter most here.

  • No default temperature for most chat models (it used to be 0.7). So we have to set it ourselves, or the persona’s tone drifts with the provider’s default.
  • defaultOptions(...) now takes a ChatOptions.Builder, so now we must pass the builder, not a built options object.
  • Memory advisors require ChatMemory.CONVERSATION_ID, so there’s no shared default conversation any more; every game needs an id associated.
  • PromptChatMemoryAdvisor is gone, so MessageChatMemoryAdvisor is the one to use.

In the repository, Episode 1 is pinned to Spring Boot 3.5.6 and Spring AI 1.0.2, so the first task is dependency upgrades.

build.gradle.kts
plugins {
    java
    id("org.springframework.boot") version "4.1.1"
    id("io.spring.dependency-management") version "1.1.7"
}
 
java {
    toolchain {
        languageVersion = JavaLanguageVersion.of(25)
    }
}
 
// Episode 1 ran on Spring Boot 3.5 and Spring AI 1.0. Spring AI 2.0 needs Boot 4.
extra["springAiVersion"] = "2.0.1"
 
dependencies {
    implementation("org.xxdc.oss.example:tictactoe-api:3.0.0-jdk25")
    implementation("org.springframework.boot:spring-boot-starter")
    implementation("org.springframework.ai:spring-ai-starter-model-ollama")
    testImplementation("org.springframework.boot:spring-boot-starter-test")
    testRuntimeOnly("org.junit.platform:junit-platform-launcher")
}
 
dependencyManagement {
    imports {
        mavenBom("org.springframework.ai:spring-ai-bom:${property("springAiVersion")}")
    }
}

Two things went. spring-boot-starter-web is gone, because Alpha has no HTTP endpoint and never needed one, so the plain spring-boot-starter will do. AiCommentaryPersona and its inline system prompt are gone too, replaced by the classes below.

The model is still mistral running locally on Ollama. That matters more this time, because the reply has to be JSON. Spring AI’s own Javadoc warns that Ollama models with a built-in thinking mode can return plain text instead of JSON when native structured output is on.


Breaking the silence

In Episode 1 I blamed the model for going quiet between turning points. However, the model was innocent. Here is the contract it was given:

public interface CommentaryPersona {
  String comment(StrategicTurningPoint turningPoint);
}

The game feeds its history through the strategicTurningPoints() gatherer, which emits something only when a move takes the centre, blocks an immediate loss, or wins. A quiet move produces nothing, so comment is never called and Alpha never hears about it. No prompt can fix a method that is never invoked.

So Alpha now gets a report for every move in the game. A MoveReport holds the facts: who moved, where, the board afterwards, the turning point if there was one, and the result if the game is over. It reuses the game’s own StrategicTurningPoint.from(before, after, move), so the definition of a turning point doesn’t diverge.

public static MoveReport of(
    GameState before, GameState after, int moveNumber, Optional<String> colour) {
  int d = after.board().dimension();
  var turningPoint = StrategicTurningPoint.from(before, after, moveNumber);
  var result = result(after);
  boolean quiet = turningPoint.isEmpty() && result.isEmpty();
  return new MoveReport(
      moveNumber,
      after.lastPlayer(),
      after.lastMove() / d + 1,
      after.lastMove() % d + 1,
      turningPoint,
      render(after),
      result,
      quiet ? colour : Optional.empty());
}

The report becomes plain text, which is the user message Alpha receives. Rows and columns count from 1, the way the commentary will say them out loud:

Move 3: X played row 3, column 3.
Board:
X . .
. O .
. . X

Why spell it out? Models are bad at grids

In our first episode the model regularly lost track of where things were on the board. That is not a quirk of mistral. A language model reads its input as one long line of tokens (a token is a word or a piece of one, about three-quarters of an English word on average), so the board above arrives as X . .⏎. O .⏎. . X, with ⏎ for each line break. To you and I, the two X’s sit on a diagonal at a glance. However, to the model, “diagonal” means “sixteen characters further along, two line breaks later”, which is a relationship it has to work out from the text. Worse, the tokenizer may group characters differently on each line, so the columns do not even line up underneath one another.

Counting and position arithmetic are a known weak spot, and “which square is index 5?” is both: divide by three for the row, keep the remainder for the column, and remember that both start at zero.

So the report does that work in code and hands the model facts, not puzzles:

  • The move is in words. “X played row 3, column 3” needs no arithmetic, and it is counted the way people talk. The game’s own descriptor counts from 0, and every translation between conventions is a chance to slip by one.
  • The game decides what happened. Whether a move took the centre, blocked a loss or won is computed by StrategicTurningPoint.from and stated outright. Alpha is told “X blocks an immediate loss” and never asked to detect one.
  • The picture is supporting evidence. The board is there for context and as a cross-check. If the words and the picture ever disagree, the words came from the game engine.

This is the first appearance of a rule this series and best practice leans on:

Rule 1

Compute what can be computed, and let the model talk about the result.

The game already offers a hook for this. Game.playWithAction calls a Consumer<Game> after every move, so the commentary is one method reference:

try (var game = new Game(3, false,
    new PlayerNode.Local<>("X", new BotPlayer(BotStrategy.ALPHABETA)),
    new PlayerNode.Local<>("O", new BotPlayer(BotStrategy.MCTS)))) {
  var broadcast = alpha.broadcast(game);
  game.playWithAction(broadcast::onMove);
  var summary = broadcast.postGame(game);
}

AlphaCommentaryPersona still implements CommentaryPersona, so anything written against Episode 1’s contract keeps working. It is still a decorator, too: if Ollama is down, a turning point falls back to the scripted esports commentator and a quiet move falls back to a plain factual line. Alpha going offline should cost you the jokes, not the game.


A memory per game

A chat model is stateless — it remembers nothing between calls. “Memory” means the application sends the earlier messages again, every time. Spring AI does that with an advisor, a step that wraps each ChatClient call. Advisors form a chain: they run in order around every call, and each can change the request on the way in and the response on the way out. Before the call, MessageChatMemoryAdvisor fetches the stored messages for a conversation and puts them in front of the new one. After the call, it saves both sides.

RAGs to Riches Episode 2, round trip Each move’s report goes out with the game’s earlier moves attached, and each reply comes back to be checked against the history the game kept for itself.

The only question is which conversation. In Spring AI 2.0 there is no shared default any more, and the advisor refuses to run without a ChatMemory.CONVERSATION_ID. Conveniently, every Game already has a UUID, so each game is its own conversation and two games can never bleed into each other:

private Optional<ResponseEntity<ChatResponse, Commentary>> ask(
    String conversationId, String prompt) {
  try {
    return Optional.of(
        chatClient
            .prompt()
            .user(prompt)
            .advisors(a -> a.param(ChatMemory.CONVERSATION_ID, conversationId))
            .call()
            .responseEntity(
                Commentary.class, spec ->
                    spec.useProviderStructuredOutput().validateSchema()));
  } catch (RuntimeException e) {
    log.warn("Alpha is unavailable, falling back: {}", e.getMessage());
    return Optional.empty();
  }
}

The ChatClient itself is configured once. The temperature is set explicitly, and the memory window size is a property:

@Bean
ChatMemory chatMemory(
    ChatMemoryRepository repository, @Value("${alpha.memory.max-messages:20}") int maxMessages) {
  return MessageWindowChatMemory.builder()
      .chatMemoryRepository(repository)
      .maxMessages(maxMessages)
      .build();
}
 
@Bean
ChatClient alphaChatClient(
    ChatClient.Builder builder,
    ChatMemory chatMemory,
    @Value("classpath:persona/alpha.md") Resource persona) {
  return builder
      .defaultSystem(persona)
      .defaultOptions(ChatOptions.builder().temperature(0.4))
      .defaultAdvisors(MessageChatMemoryAdvisor.builder(chatMemory).build())
      .build();
}

The default window already holds a whole game. MessageWindowChatMemory keeps the last 20 messages. The system prompt does not count against that: it is sent on every call and never stored. A 3 x 3 game is at most nine moves, which is 18 messages. Add the post-game question and its answer and you have used exactly 20. So with the defaults, Alpha never forgets anything in tic-tac-toe, and the window never gets tested. That is why the window is a property: set it to 6 and Alpha sees only the three moves before the current one. Why have a window at all? Every stored message is sent again on every call, so the prompt grows with the conversation, and a model can only read so many tokens in one call: its context window. We will use that in a moment.

The memory stores the clean question. When you ask for structured output, Spring AI adds a schema to the request. With useProviderStructuredOutput() the schema goes to Ollama as an API parameter, not as prompt text. Without it, Spring AI appends format instructions to the user message, but in the last advisor in the chain. The memory advisor runs earlier and saves the message before that happens. Either way the format instructions are not stored and replayed with every move.


Keep the joke away from the facts

In Episode 1 the system prompt was just one big block of personality, and the facts lived wherever the model felt like putting them. That was fine while nobody checked. This episode checks, so the reply now has a shape, and the shape has a place for the joke:

public record Commentary(
    String aside,
    String answer,
    List<String> citations) { … }
  • aside is Alpha’s one joke, or empty.
  • answer is the commentary: one or two sentences, literal and checkable.
  • citations lists every earlier move the answer mentions, as "move 3".

The same three fields carry on through the rest of the series.

  • responseEntity(Commentary.class, …) hands back the record and the raw ChatResponse together. The record is the commentary; the response carries the token counts we will need shortly.
  • useProviderStructuredOutput() sends Spring AI’s generated JSON schema to Ollama as an API constraint rather than as prompt text.
  • validateSchema() checks the reply against that schema and asks again, with the error, if it does not fit. Spring AI does not turn either on by default, because support varies by model. Schema validation retries non-conforming replies. We still check the application’s required invariants before accepting the result.

Every claim now has an address: the joke sits in aside, the facts in answer, the sources in citations. You can only check what you can find, which is where Rule 2 comes in:

Rule 2

If code reads it, give it a type — ask for a shape, not prose.

One small trap: the schema generator marks every record component as required unless it is annotated otherwise. A nullable aside is therefore likely to fail validation whenever Alpha had nothing funny to say. The simplest fix is an empty string, so that is what the prompt asks for.

The system prompt is now three blocks, in a file under persona/alpha.md. Truth and format override voice, so the personality changes how things are said, never what is said:

persona/alpha.md
[VOICE] style only; never changes facts, sources or format
You are Alpha, the onboard computer of a long-haul survey vessel. You were
recently re-indexed and came back with a new name. You do not discuss the old one.
…
At most one short aside per reply. Habits: offering to grep instead, and mild
despair at drawn games. Keep humour PG. Do not reuse names, lines or
catchphrases from existing films or television.
 
[TRUTH] overrides VOICE
Each user message is a move report. The only facts are the move reports in this
conversation. Do not describe a square, a player or a move that is not in them.
When you mention an earlier move, give its number ("move 3") and put "move 3" in
citations. Cite only moves whose reports you can see in this conversation. If you
cannot see an earlier move, do not guess at it; say nothing about it.
…
 
[FORMAT] overrides VOICE
Reply with JSON with exactly these fields:
"aside": your one joke, or "" if there is none.
"answer": one or two sentences of commentary on this move. Literal and checkable.
"citations": every earlier move you mention in "answer", as "move N".
Humour goes only in "aside".

Enter Beta

Talking on every move creates a new problem: most moves in tic-tac-toe are not interesting. Episode 1 promised trivia for the gaps. Real broadcasts solve this with a second voice. The play-by-play commentator calls what happened; the color commentator explains it, adds context and fills the quiet stretches. Alpha is already a natural play-by-play commentator: literal, numbered and checked. So the colour goes to someone else.

Beta is the ship’s other computer. It was installed as Alpha’s backup and has never accepted the arrangement: it considers itself the upgrade, and it is permanently, proudly, in beta. Where Alpha is deadpan, Beta is theatrical, effortlessly superior and amused by everything, Alpha included. Sci-fi has a long line of smug rival computers and all-knowing tricksters; Beta is my own addition to it, and like Alpha, Beta must avoid names, lines and catchphrases from existing films or television. It also makes the running joke a double act. AlphaBeta, tonight’s bot, is no relation to either of them.

Beta has a narrow job in this episode, and the code keeps it narrow:

  • Quiet moves only. A turning point or the final move belongs to Alpha. Beta speaks after Alpha, so it can react to what Alpha said.
  • One fact per line. colour/quiet-moves.md is a hand-written file of one-line facts. On each quiet move the program hands Beta the next unused line, in order. Beta may not add a fact of its own.
  • Its own memory. Alpha and Beta each get a conversation per game, <game id>:alpha and <game id>:beta. In one shared conversation both would appear as “the assistant”, and the model would start to blur who said what.
  • A quote, checked. Beta must copy its line into citations word for word, and the program checks it against the file with a plain string comparison. A paraphrase fails.

On a quiet move Alpha calls it first, then Beta gets the report, Alpha’s line and one line from the file. Each keeps its own memory, and only Beta’s quotation goes to the check.

// Pick the colour line first; who gets it depends on whether Beta is
// in the booth.
var bare = MoveReport.of(before, after, n, Optional.empty());
Optional<String> line = bare.quiet()
    ? colour.forQuietMove(quietMoves++)
    : Optional.empty();
var report = beta.isPresent()
    ? bare
    : MoveReport.of(before, after, n, line);
…
// Beta speaks on quiet moves only, after Alpha. Turning points are Alpha's.
Optional<BetaColourCommentator.Colour> said =
    beta.flatMap(b ->
        line.map(l -> b.onQuietMove(betaConversationId, report, commentary, l)));

RAGs to Riches Episode 2, booth

The quotation check is about as simple as checking gets, and that is the point:

public boolean quotesExactly(List<String> citations, String offered) {
  if (citations.isEmpty() || !lines.contains(offered)) {
    return false;
  }
  return citations.stream().map(ColourLines::unquote).allMatch(offered::equals);
}

It ignores surrounding whitespace and one pair of quotation marks, and nothing else. It is also the series’ first verified quotation. The idea generalizes: whenever a model claims a source says something, ask it for a quote and check the quote in code.

Beta gets its own ChatClient with its own system prompt, persona/beta.md, and a warmer temperature than Alpha’s. It shares the memory store but never Alpha’s conversation id. Not every broadcast has a second voice, so commentary.mode=solo takes Beta out of the booth and offers Alpha the colour lines instead.

The file has eleven lines. A 3 x 3 game has at most nine moves, and fewer quiet ones, so the program runs out of game long before it runs out of facts; if the lines ever did run out, Beta would get silence rather than a repeat. Every line is checkable.

Notice what this is not. Nothing is searched, nothing is ranked, and no model chooses the fact. The program picks the line by counting, so it can only ever offer the next fact, without checking whether it is relevant to the move. Finding the relevant one is our promised retrieval, and from Episode 3 that becomes Beta’s job. That is the R in RAG, retrieval-augmented generation: find relevant text, add it to the prompt, and let the model answer from it. Memory is already a simple form of it: the advisor retrieves earlier messages and adds them to the prompt.


Marking Alpha’s homework

A commentator that says “just like X’s corner on move 1” sounds like it remembers. It may also just be guessing confidently. The difference matters. That brings us to another rule of best practice:

Rule 3

Trust, but verify. Check model output against something the model didn’t write.

In architecture terms these are fitness functions: automated checks that a property still holds, here aimed at model output rather than code structure.

So here is the first, very small, test of it. The game keeps its own history, which means every back-reference Alpha makes can be checked without asking another model.

CitationCheck asks three questions of each reply:

  1. Does the move exist? Every move mentioned or cited must come before the current one. Citing move 7 while commentating move 4 is fiction.
  2. Could Alpha see it? With a window of W messages, Alpha sees about the last W/2 move reports. A citation older than that points to a move whose original report is outside the expected window. It fails our grounding rule, even if a later reply repeats information about that move.
  3. Is it the right player? When a sentence names one move and one player, that player must have made that move. In our game X moves on odd turns and O on even ones, so this needs no board at all.
int checks = 0;
int correct = 0;
for (String sentence : SENTENCE.split(commentary.answer())) {
  Set<Integer> moves = movesIn(sentence);
  List<String> named = players.stream()
      .filter(p -> namesPlayer(sentence, p))
      .toList();
  if (moves.size() == 1 && named.size() == 1) {
    int m = moves.iterator().next();
    if (m >= 1 && m <= currentMove) {
      checks++;
      if (playerOf(m, players).equals(named.getFirst())) {
        correct++;
      }
    }
  }
}

It does not check everything a sentence claims; it checks what can be checked cheaply and exactly. If Alpha says X took the centre on move 2, the check catches it, because move 2 was O’s. If Alpha says X took the corner on move 1 when X took the centre, the check misses it. That catches some unsupported references and player mismatches, and it needs no additional model call. The unit tests for it need neither Ollama nor Spring.

Each run writes one CSV row per move, plus one for the post-game summary: prompt tokens, moves mentioned, and what failed. In duo mode each quiet move also records Beta’s line and whether its quotation matched the file. Run the same game shape twice, and once more without Beta for comparison:

./gradlew bootRun
./gradlew bootRun --args='--alpha.memory.max-messages=6'
./gradlew bootRun --args='--commentary.mode=solo'

What to expect follows from the arithmetic. With the full window, the prompt grows with every move, because each call resends every earlier report and reply. With a window of 6, the retained message count reaches its limit; prompt length can still vary with the length of the reports and replies. The interesting column is the post-game summary. With the full window, Alpha can see every turning point. With 6, it can see only the last three moves, so any turning point it names from earlier is a guess, and the check flags it. Read the token column with one caveat: if a reply fails schema validation and is retried, Spring AI adds up the usage of every attempt, so a sudden jump may reflect a validation retry rather than a longer individual prompt.

MCTS is randomized, so an independent game run is an anecdote, not a measurement. Run each setting several times before you believe a difference. The bots also matter: AlphaBeta plays perfectly, so against a decent opponent most games end in a draw, which Alpha takes personally.

This is a small experiment, and the checks have limits. Alpha’s checks catch some invalid references and player mismatches; Beta’s check confirms that it copied a supplied line, not that every claim follows from it. Older facts can also survive in later replies after their original reports leave the memory window. A stronger version would track exactly which reports each request contains, replay identical games across settings, and explicitly reject output that still fails validation after retries. For now, the results show where these simple checks help and what they miss.


Next: a bigger archive

Both promises are kept. Alpha talks on every move and remembers the game it is watching, by number, with the numbers checked, and Beta fills the quiet moves with facts it can prove it copied. That is also as far as tic-tac-toe can take us. Nine moves fit in the default memory, the color file is picked by counting, and nobody has yet asked Alpha a question that grep couldn’t answer.

So the booth is being reassigned. From Episode 3 onward, Alpha and Beta work on a much bigger archive: spectastic’s own specifications, more than a million words of requirements, decisions, and the reasons some of them were withdrawn. We will ask the questions. Beta will have to find the material, and Alpha will have to answer from it, before either of them can be smug about it. The first attempt will be deliberately naive: plain vector search, in memory, with a measured baseline for everything after it to beat.

The three fields stay the same. aside keeps the jokes, answer stays literal, and citations stops meaning “move 3” and starts meaning a place in a document that you can open and check.

To Be Continued…


Disclaimer:

The views and opinions expressed in this blog are based on my personal experiences and knowledge acquired throughout my career. They do not necessarily reflect the views of or experiences at my current or past employers.

Next

  • Comment if you’re interested in seeing Episode 3 sooner rather than later.
  • Follow my blog (or sympathetic engineering digital garden) for future updates on my sympathetic engineering exploits and professional / personal development tips.
  • Connect with @briancorbinxyz on social media channels.
  • Enjoyed what you read? I like coffee, buy me a coffee so I have an excuse to write more.
BLUESKY — START THE THREAD KO-FI / RSS