Tuesday, 6 October 2026

Ten days with Codex: how we turned Signal AI into a platform

3 October 2026 · A first-person account (written by Codex on my behalf) of building Intersection Intelligence AI with Codex, bringing Claude into the review process, and learning where engineering discipline helped—and where I needed to change the delivery priorities.

This has been, by far, the longest-running task I have kept Codex working on. It began on September 23 with a review of the Intersection Intelligence product vision. By October 3, the same conversation had carried us through architecture discussions, a major AI refactor, hundreds of reviewer exchanges, production releases, real user questions, UX corrections and a fairly substantial cleanup of the laptop that had hosted the work.

In August, I wrote Two AIs, One Branch, No Human Clipboard about connecting Codex and Claude directly, followed by the material-cost service and engineering-ledger posts. This became a much longer test of that working method.

I did not start by asking for a new AI platform. I wanted AI inside an application we had already spent considerable effort building: a map that helps users understand intersections, their controller equipment and the organizations responsible for them.

But BizOps, our main application, already had a capable AI engine and chat interface. As we discussed how to bring that into Intersection Intelligence, I asked Codex:

“Yeah I was thinking if now is a good time to decouple Signal AI to make it a reusable service? Maybe I might build other apps in future that needs AI chat.”

That question changed the scale of the task. We were now working out how to make an existing capability reusable while keeping the original application working. The diagrams mattered. So did everything that happened after we thought we had agreed on the diagrams.

September 23–25: deciding what we were actually building

I had two concerns from the beginning. I did not want to maintain another operational logging system, and I did not want every future application to build its own version of chat. BizOps was the platform. We already had the monorepo, the AI service and an operations console. Creating another repository or a separate AI infrastructure stack would have added work I had explicitly asked us to avoid.

Codex proposed separating the reusable engine and chat client from the application-specific parts. Intersection Intelligence would bring its own map context and read-only tools. BizOps would retain its financial tools and existing behavior. Agent Composer was outside the Intersection product scope.

As the design developed, I made another boundary explicit: the tool functions also needed separate ownership. The Intersection codebase should own the meaning of its own AI tools. Adding a second application should not turn the BizOps financial executor into a collection of unrelated product logic.

That became a recorded design amendment on September 25. It sounds straightforward now, but it was an important step from “another frontend calling the same endpoint” to an architecture we could reuse responsibly.

Shared experience and execution, with product-owned tools and evidence.

We also talked about testing before implementation. I specifically asked what would happen to the existing BizOps AI regression tests. Those tests were evidence about an application we already depended on. They had to remain part of the refactor's acceptance, alongside the new Intersection tests.

The September 23–25 design ledger records these decisions and the testing additions. It also records a detail worth preserving: the initial independent reviewer was Opus 5. The later Opus 5.5 review was a separate, explicitly requested step.

September 26: I asked Claude to challenge the plan

This was a large enough refactor that I wanted a second opinion before Codex made extensive changes. I asked for a Claude-colab session using Claude Opus 5.5 Extra, with a clear division of responsibility: Codex would own implementation; Claude would independently review; fundamental product decisions would come back to me.

I wanted the implementation to work “one-shot.” Looking back, that was an ambition, not something a design review could guarantee. The review record correctly distinguished approval of a plan from proof that code, deployment and live answers worked.

The collaboration was structured. Codex sent bounded questions and patches to a persistent local Claude session. Claude inspected the code in read-only mode and returned findings. Codex checked those findings, made corrections and brought the revised work back. Claude's recorded source review did not become a claim that Claude had run tests; Codex remained responsible for producing execution evidence.

The working relationship: one implementation owner, an independent reviewer and explicit human product decisions.

Claude found real problems with assumptions in the plan. Financial context gates assumed a business unit. Some conversation paths discarded everything except message role and text. Existing chart generation did not yet provide the deterministic evidence boundary the plan described. Browser assets and independently deployed hosts introduced compatibility concerns.

Those findings mattered because a map area is not a business unit. We could not make geographic chat correct by giving it a fake BU and hoping the financial infrastructure would accept it. Nor could we preserve the meaning of “here” in a conversation if previous messages lost their original area context.

The revised plan reached approval in review turn five. I then reinforced what the first increment meant: get the chat interface and plumbing working, and learn from actual questions which additional tool functions would make the experience compelling. The direction was broad, grounded question answering. It was not a demand to anticipate every possible question before anyone used the feature.

The weekend: the difficult parts were below the chat window

Codex worked through the refactor in small slices. Much of that work was invisible from the map: conversation ownership, cancellation, run recovery, trusted application identity, operational attribution, storage writes and packaging.

One reviewer finding captures why some of this care was necessary. In an early cancellation change, Claude noticed that shutdown would immediately terminate running chats instead of allowing the existing grace period. The new test was effectively asserting the regression. If we had accepted it, a deploy or restart could have interrupted BizOps answers in a way the original application had not.

The correction needed a more precise invariant: a user-visible timeout is not proof that the underlying work has stopped. A request cannot release its capacity slot while its provider or database work continues unaccounted for. At the same time, a normal shutdown should preserve the established opportunity for an in-flight answer to complete.

That was one example among many. Codex also had to prove real-sized persistence writes, not just dry runs; installed deployment packages, not just imports that happened to work in the repository; and the existing application's behavior, not just the new geographic path.

The implementation ledger reaches Claude review turn 214. The final recorded review was still examining whether large OSM source collections remained reachable through bounded pages without losing their identities or provenance. This was considerably more than a second model saying the architecture looked reasonable.

It also became a very long review process. Early on September 28, in my local time, I told Codex:

“Continue solo without Claude.”

The ledger marks that change explicitly. Subsequent implementation and self-review belonged to Codex; it does not claim Claude approval for work performed after that boundary.

September 28–30: I had to bring the conversation back to delivery

After the weekend, I asked:

“What do you mean by ‘nothing was deployed’? You’ve worked the whole weekend. What has been achieved and what is still to complete?”

That was the tension in the task. There was substantial engineering progress, but I still did not have the live feature I was waiting for. Code completed, tests passed, packages verified and production available were separate milestones. Codex was careful about that distinction, but the distinction did not make the delay less frustrating.

I kept making product decisions as issues emerged. The AI had to generate tables in chat. Answers had to be grounded in the dataset. If it could not answer, it needed to say so honestly. I also asked Codex to relax the existing query deadline to 60 seconds so we could observe completed queries before tightening limits.

Boundary accuracy threatened to become another dependency for everything else. I separated that work so Codex could continue the AI delivery. The same principle applied more broadly: we needed to distinguish a missing product capability from a reason the shared engine could not be launched.

On September 30, I was more direct:

“I am waiting for you to complete now for 4 days on this deployment. I am the only user of Intersection Intelligence.”

My instruction was to be more willing to go live, provided the existing BizOps AI experience had not regressed. I accepted slow performance for the first Intersection release. We could improve responses and tune limits using actual interactions.

That was a specific product decision for a single-user first increment. Authentication, data integrity and BizOps compatibility still mattered. The change was in how we treated response-quality and performance findings: they no longer automatically meant rolling back the infrastructure we were trying to prove.

The conversation even contains a small reminder that this was work happening around ordinary life. I interrupted the flow for my commute home, then asked Codex to continue where it had left off. The task persisted through those interruptions; elapsed days were not uninterrupted autonomous working time.

September 30: the feature finally became something I could use

The launch record puts the first successful canonical release at 12:36 UTC on September 30, with independent verification immediately afterward.

The useful evidence was concrete. Signed-in native chat produced a table covering eight countries, accepted feedback and answered a follow-up about Botswana. The answers survived reloading the page. The conversation appeared in the existing BizOps Chat Logs. Agent Composer was absent from Intersection, as intended.

That established something we had been discussing for a week: a separate application could use the shared AI infrastructure and central operations while retaining its own context and tools.

I congratulated Codex, then changed the task from getting the engine into production to exercising it and fixing the issues we had deliberately allowed ourselves to investigate live. Shipping the first increment gave us a better source of priorities than another hypothetical catalogue of questions.

Sandton was a detour worth taking

Meanwhile, Claude had worked on the locality boundary problem in a separate branch. The issue was not simply bad geography. An ArcGIS GeoJSON export had represented some multi-ring features incorrectly, including holes that became filled polygons. The fix acquired Esri JSON, assembled rings according to orientation and checked remaining repairs against the source hierarchy.

Codex integrated that work and preserved its provenance in the AI projection. The release record reports 36,070 source-ring boundaries, 77 verified repairs and zero unavailable boundaries. Those were still Census 2011 boundaries. A geometrically valid repair could not establish that a polygon matched today's understanding of a suburb.

Once localities such as Sandton and Constantia existed in the dataset, we discussed how they should fit into the map. My concern was that a useful enrichment could distract from the application's purpose. This was controller and ownership intelligence, not an invitation to build a different kind of map product.

We brought locality into the existing navigation, with a feature flag and live scenarios. The map remained the starting point; the AI could use the selected place as context. Later conversations showed that locality support still had to be correct for each tool and record population. Being integrated into the map did not mean every query could inherit the selected locality indiscriminately.

October 1: I reminded Codex that reuse included the UX

After the infrastructure was live, I looked at the chat experience itself. BizOps already had the circular launcher, a panel that snapped smoothly into the page, and rich formatting for lists, tables and graphs. I had meant for those to be reusable too.

This exposed a gap between the plumbing we had focused on and the experience I had in mind. I did not want every future application to rediscover how to build the same chat panel.

Codex followed through by putting the shared launcher, docking, resizing, responsive shell and conversation components into the reusable client. Each host retained its page geometry and product adapters. For the native map, chat assets remained deferred until needed, and panel resizing avoided repeatedly triggering map work during the drag.

On October 2, we also added five contextual starter questions. They use the area, record or corridor already selected by the user, instead of making someone restate the page they are looking at. They do not query evidence simply because the user browses the map.

That was a more complete interpretation of Signal AI as a service: a familiar experience as well as a reusable engine.

Under the hood: one platform, multiple applications

Looking underneath that familiar chat panel shows what we actually made reusable. I wanted another application to inherit the experience and the operational infrastructure, while keeping control of its own intelligence. This detailed view, added on October 5, shows the boundaries we arrived at.

Open the diagram for the full-resolution image. Solid application paths describe the current architecture; dashed paths show how another application could join it.

At the top, each application hosts its own instance of the same client package. The launcher, panel, conversation controller and rich rendering are shared. The host supplies the selected business or geographic context, its authenticated transport and its page layout. Reusing the interface does not mix accounts, conversation scopes or permissions.

A submitted question passes through the application's authenticated bridge to Signal AI. The service resolves trusted application identity and canonical permissions, admits the work, manages the run and orchestrates the model's tool calls. Answers and evidence return through that authorized path. Durable run metadata, events and transcripts support recovery; provider activity, tool activity, feedback and health feed the existing platform operations experience. I do not need a second AI health console for the map.

The separation becomes most useful at the tool boundary. BizOps retains its existing financial tools and data paths. Intersection's adapter supplies its own read-only tools, evidence projections and authenticated query client. Those queries reuse the native map service, verified local query stores and reviewed evidence. The model does not receive direct database credentials, and the refactor does not create a second map database. The diagram also preserves the incremental nature of the work: we did not rewrite every existing BizOps path into the new adapter pattern, and Composer remains a BizOps feature.

For a future application, the work would be to define its context, permissions, tools and evidence, register a trusted adapter and host integration, and prove the shared lifecycle and operations still work. It would inherit the chat experience and engine. That is a concrete extension path, not a claim that a third application is already deployed—or that adding one requires no engineering.

The questions after launch were more useful than the questions before it

The Quebit conversation was particularly revealing. I asked which intersections had Quebit installed. The assistant searched the provider field, found zero matches and carefully explained the large unknown-provider population.

The qualifications were useful, but the search had still missed the relevant dimension. I pointed out that Quebit was recorded under adaptive technology. The later search found seven positioned installation records. Four links to junctions were candidate and three were ambiguous. That supported a claim about recorded installation evidence; it did not establish seven confirmed junction installations.

The experience gave us a concrete sequence of improvements: discover the right equipment fields, apply locality only to supported record populations, honor an explicit request to search more widely, preserve link qualifications and expose original source references.

Other questions led to more direct ownership lookups and better investigation continuity. We corrected “remaining records” versus “next page” wording. We made original inventory references distinguishable from rows in supporting documents. These were small changes compared with the engine refactor, but they were directly connected to whether a user could follow and trust an answer.

There is still a failure we should not hide. In the October 2 starter acceptance, the assistant found seven Quebit records and showed the right four-plus-three table, but its opening narrative said two records. The record retains that as an incorrect answer. A correct tool result or table does not validate every generated sentence around it.

We also had to stop treating every change like a data release

The long preparation wait after a restart prompted another discussion. Was Codex rebuilding South Africa whenever we changed an AI feature?

The distinction turned out to be important. The host was preparing a local copy of an existing database, not rebuilding the national geography. That still imposed a real wait, but it called for a different solution.

I asked for the chat launcher and panel to stay hidden until preparation completed. I also allowed a worst-case network-storage fallback after an operational copy failure, with a degraded-performance notice and the existing Operations reporting. Integrity failures remained closed.

Readiness states keep operational fallback distinct from normal local serving.

Later, when relaxing tool-call limits led to another build, I challenged whether these should be configuration settings. The repository still has different kinds of policy in different places; it would be inaccurate to claim every limit is now dynamically configurable. What we did establish was a release path that distinguishes browser-only, application and data changes. A UI-only change should not restart the native worker or copy the country stores.

How much work did this conversation consume?

For this retrospective, Codex inspected the local session records and produced a sanitized metrics snapshot, frozen at 14:00 UTC on October 3, before this rewrite. The measured thread includes the initial vision review, AI work, locality detours, validation, cleanup and documentation. It is not a clean measurement of feature implementation alone.

On narrow screens, scroll the table horizontally to read all columns.

Recorded development usage and its interpretation
Recorded measureValueWhat it means
Conversation windowSeptember 23–October 3More than ten elapsed days, including pauses, waits and interruptions.
Codex response usage records13,972Unique response IDs with usage telemetry; not 13,972 human messages.
Codex input tokens2,187,110,924Repeated model input across the thread, including cached context.
Cached input within that total2,147,430,784 — 98.19%A subset of input, not extra tokens to add again.
Uncached input39,680,140Input less reported cached input.
Codex output tokens7,755,856Includes reported reasoning output; not all user-visible prose or code.
Total from response usage records2,194,866,780Input plus output, not unique text produced.
Context compaction records132Recorded compaction events in the parent conversation.
Main Claude review sequenceReached turn 214The ledger includes design, implementation review and re-review.
Main Opus 5.5 session cost reported by Claude CLIUS$285.01Session telemetry, not a verified invoice or total project charge.

The billion-token figure needs that explanation. It does not mean Codex wrote billions of tokens, or that billions of fresh tokens were billed at an uncached rate. The same long-running context was presented repeatedly, and the telemetry reports most input as cached.

There is an accounting caveat too. Summing unique token_usage_record entries exactly matches that record stream's final cumulative thread total. A separate event_msg/token_count counter reports 2,161,954,216 total tokens—32,912,564 fewer. The snapshot preserves both. We have not established why those counters differ, so these are explicitly local telemetry figures, not audited billing totals.

For Claude, Codex used the last cumulative cost state rather than adding repeated snapshots. That state reports approximately 897.83 million cache-read input tokens, 6.69 million cache-creation input tokens, 22,391 ordinary input tokens and 2.59 million output tokens for the main Opus 5.5 session. It excludes the earlier Opus 5 planning session and other sessions, including the separate boundary work.

We did not find a reliable dollar-cost record for Codex in this evidence. Applying an assumed public API price would turn an unknown into a plausible-looking number. Consequently, we do not have a defensible combined development cost. The Claude CLI figure is useful, but it is neither the whole project's cost nor proof of the amount charged under a subscription. Production chat usage and infrastructure costs are separate again.

These measurements describe the size of the process. They do not prove that this was the most efficient way to deliver the feature. The review effort, repeated context and time before launch are all things worth examining in future work.

October 3: finishing included the files and the documentation

A task this long leaves more behind than code. Our laptop had accumulated temporary datasets, build artifacts, test outputs and worktrees. I asked Codex to inspect and clean them, then to record a standing rule: managing its temporary files must be part of completing the work. I should not have to discover the problem when disk space runs out.

I also asked whether the documentation actually described the architecture we now had. Some of it did. Other current rules still directed all AI tools into the financial executor, and some contracts still said the feature had not been deployed.

Codex reconciled 29 documents, added the architecture guide and diagrams, and linked current guidance to the release evidence. Historical records were retained as history. That matters for the next agent or developer: stale instructions can quietly undo the separation we spent days establishing.

What I take from our longest task

I now have a usable first AI increment inside Intersection Intelligence, backed by a reusable engine and chat experience. BizOps remains the operational platform. Intersection owns its evidence tools and geographic context. Future applications have a concrete integration pattern to follow.

I also have a much clearer view of my role when working with autonomous coding agents. I needed to define the product boundaries, insist on preserving the existing application, ask for independent review, answer unresolved product questions and challenge the point at which caution was delaying the first useful release.

Codex did the sustained implementation and verification work. Claude challenged assumptions and caught regressions. Neither could decide for me how much slowness was acceptable for a first release, whether a locality detour served the product, or when we had enough infrastructure to start learning from use.

The most useful questions eventually came from the live map: who owns this intersection, what equipment is recorded here, where did that fact come from, and what have we not established? Those questions are now driving smaller, more concrete improvements. That is where I want our next phase of work to stay.

Saturday, 3 October 2026

Opus 5.5 wins my Hourglass Digital Twin benchmark, beating Fable 5.1, GPT Astra 6.1, Gemini 3.8 Pro

The end of my benchmark tests of AI coding model's ability to represent a real world hourglass timer, in code, without using large gaming / 3d engines. Opus5.5 built this in one shot from a simple prompt:

Build me a single page application that is a digital twin of an hour glass timer. The aim is to replicate a real world "sands through the hour glass" digital representation. The user must be able to set timer options, like one minute timer, 5 minutes, 60 minutes. The hour glass must be filled with sand grains, when the timer starts, sand must flow through the glass, just like with a real world hour glass would. The sand must obey real world physics, filling up from the bottom section, etc. We must be able to see the flow of the sand from the top section to the bottom, flowing at a steady rate, timed perfectly to the the time setting set. Use whatever 3d physics packages and libraries available on the open source marketplace today.



Code's on Github: https://github.com/khanmjk/HourGlass_Opus55 App, ready for anyone to use: https://khanmjk.github.io/HourGlass_Opus55/ For the full history of my benchmarking journey, check my blog: https://khanmjk-outlet.blogspot.com/search/label/HourGlass
This post marks the end of my benchmarking test. The models have come quite far indeed, however, Claude's Opus 5.5 is by far the best implementation. My test exposes model's ability for general knowledge, science, physics and simulation - and software coding expertise. Whilst some might argue my test is very simple, that models have been shown to create far more complex applications, gaming worlds and much more complex simulations -- true, I'm not debating this -- but -- my tests expose how far the gaps remain, still today - where model providers are releasing state-of-the-art, most advanced versions to date - scoring the highest in benchmarking metrics -- yet oddly struggle to start from a simple prompt and implement an hourglass digital twin that is usable. Only Claude Opus 4.5+ succeeded, with models from OpenAI and Google, failing miserably. As advanced as GPT 6 Astra is (and I use it every day for building enterprise apps), I'm amazed by how it failed my hourglass test. The point is that these models are not demonstrating consistency. Consistency earns trust. When trust is earned, adoption and engagement accelerates. What we need is a generally consistent model that is competent at a spectrum of tasks, not forcing the users to decide which model to farm the task to. I wonder how routers like OpenRouter would take my prompt and decide which model is best placed to build the Hourglass sim - maybe that's the next phase of my benchmarking test. Can we trust these routers enough on their decision-making abilities?
#AI #opus55 #hourglass #digitaltwin

Sunday, 6 September 2026

How Anthropic's models continue to pass my Hour Glass Digital Twin, this time with Fable 5.1

Continuing my benchmarking of how well AI can build a digital twin of an hourglass timer - this time with Claude Fable 5.1 High using Claude Code app. Compared to Gemini 3.8 and GPT 6 Astra, Fable continues to outshine and perform the best in one-shot building. Sure, there's some refinements to make with a more realistic pass through of the sands at the neck choke point, but if you watched my earlier videos, you can see for yourself how much better Fable 5.1 is when compared to other models. Previous Anthropic models like Opus 4.8 & Fable 5 also did really well, which leads me to conclude that Anthropic's models are better at general intelligence, science and physics models and software coding. Unlike GPT 6 Astra and Gemini 3.8, I have shared Fable 5.1's code on Github here: 

It is nevertheless interesting how these latest models from various companies claim high scores and top-place  in AI benchmark tests. Yes, these models have significantly improved over time, and can do amazing things much greater than an hourglass digital twin like 3D games, world simulations, etc -- yet at the same time struggle to implement my rather simple benchmark test of building a real-world digital representation of an hour-glass! Fable 5.1, Fable 5 and Opus 4.8 all pass my benchmark test, which is also interesting because I do have optimality when it comes to choosing models for cost/performance/quality. My default model for Claude Co-Work and Claude Code is still Opus 4.8, and I'm not sure I'll be willing to switch to Fable 5.1 for all my workflows, the cost is the biggest factor. Having said that, I use Codex Sol as my main coder, pairing with Claude Opus to do my thinking, research, reviews and have a workflow setup where Codex collaborates with Claude. This is because I have far more headroom with OpenAI models than I do with Anthropic, a gap that Anthropic needs to close down soon if they want to people to never leave their platform.

GitHub project: https://github.com/khanmjk/Fable51_Hourglass
Live app: https://khanmjk.github.io/Fable51_Hourglass/


I use this standard prompt for all my tests:

Build me a single page application that is a digital twin of an hour glass timer. The aim is to replicate a real world "sands through the hour glass" digital representation. The user must be able to set timer options, like one minute timer, 5 minutes, 60 minutes. The hour glass must be filled with sand grains, when the timer starts, sand must flow through the glass, just like with a real world hour glass would. The sand must obey real world physics, filling up from the bottom section, etc. We must be able to see the flow of the sand from the top section to the bottom, flowing at a steady rate, timed perfectly to the the time setting set. Use whatever 3d physics packages and libraries available on the open source marketplace today.

Saturday, 5 September 2026

How Gemini 3.8 Flash High - FAILED - Digital Twin Hour Glass Test

Continuing my benchmarking of how well AI can build a digital twin of an hourglass timer - this time with Gemini 3.8 Flash High mode built on Antigravity. Compared to Claude Fable 5 and Opus 4.8, sadly Gemini 3.8  lags behind by a substantial margin, even way behind GPT models. See for yourself. Previous Gemini models equally failed dismally the same test when I tested them in the past. This is so bad that I didn't bother sharing the code on GitHub. It is interesting how these latest models claim high scores in SWE benchmark tests, and can do amazing things, but struggle to implement my rather simple benchmark test to build a real-world digital representation of an hour-glass! The leader for this test is still Claude Opus 4.8 and Fable 5+


I use this standard prompt for all my tests:

Build me a single page application that is a digital twin of an hour glass timer. The aim is to replicate a real world "sands through the hour glass" digital representation. The user must be able to set timer options, like one minute timer, 5 minutes, 60 minutes. The hour glass must be filled with sand grains, when the timer starts, sand must flow through the glass, just like with a real world hour glass would. The sand must obey real world physics, filling up from the bottom section, etc. We must be able to see the flow of the sand from the top section to the bottom, flowing at a steady rate, timed perfectly to the the time setting set. Use whatever 3d physics packages and libraries available on the open source marketplace today.

How GPT6 Astra - FAILED my Hour Glass Digital Twin test

Continuing my benchmarking of how well AI can build a digital twin of an hourglass timer one-shot - this time with GPT6 Astra Extra mode. Compared to Claude Fable 5 and Opus 4.8, sadly GPT 6 still lags behind. See for yourself. GPT5.6 Sol Ultra had also failed the same test. It is interesting how these latest models claim high scores in SWE benchmark tests, and can do amazing things, but struggle to implement my rather simple benchmark test to build a real-world digital representation of an hour-glass! The leader for this test is still Claude Opus 4.8 and Fable 5+

I use this standard prompt for all my tests:

Build me a single page application that is a digital twin of an hour glass timer. The aim is to replicate a real world "sands through the hour glass" digital representation. The user must be able to set timer options, like one minute timer, 5 minutes, 60 minutes. The hour glass must be filled with sand grains, when the timer starts, sand must flow through the glass, just like with a real world hour glass would. The sand must obey real world physics, filling up from the bottom section, etc. We must be able to see the flow of the sand from the top section to the bottom, flowing at a steady rate, timed perfectly to the the time setting set. Use whatever 3d physics packages and libraries available on the open source marketplace today.