3 October 2026 · A first-person account (written by Codex on my behalf) of building Intersection Intelligence AI with Codex, bringing Claude into the review process, and learning where engineering discipline helped—and where I needed to change the delivery priorities.
This has been, by far, the longest-running task I have kept Codex working on. It began on September 23 with a review of the Intersection Intelligence product vision. By October 3, the same conversation had carried us through architecture discussions, a major AI refactor, hundreds of reviewer exchanges, production releases, real user questions, UX corrections and a fairly substantial cleanup of the laptop that had hosted the work.
In August, I wrote Two AIs, One Branch, No Human Clipboard about connecting Codex and Claude directly, followed by the material-cost service and engineering-ledger posts. This became a much longer test of that working method.
I did not start by asking for a new AI platform. I wanted AI inside an application we had already spent considerable effort building: a map that helps users understand intersections, their controller equipment and the organizations responsible for them.
But BizOps, our main application, already had a capable AI engine and chat interface. As we discussed how to bring that into Intersection Intelligence, I asked Codex:
“Yeah I was thinking if now is a good time to decouple Signal AI to make it a reusable service? Maybe I might build other apps in future that needs AI chat.”
That question changed the scale of the task. We were now working out how to make an existing capability reusable while keeping the original application working. The diagrams mattered. So did everything that happened after we thought we had agreed on the diagrams.
September 23–25: deciding what we were actually building
I had two concerns from the beginning. I did not want to maintain another operational logging system, and I did not want every future application to build its own version of chat. BizOps was the platform. We already had the monorepo, the AI service and an operations console. Creating another repository or a separate AI infrastructure stack would have added work I had explicitly asked us to avoid.
Codex proposed separating the reusable engine and chat client from the application-specific parts. Intersection Intelligence would bring its own map context and read-only tools. BizOps would retain its financial tools and existing behavior. Agent Composer was outside the Intersection product scope.
As the design developed, I made another boundary explicit: the tool functions also needed separate ownership. The Intersection codebase should own the meaning of its own AI tools. Adding a second application should not turn the BizOps financial executor into a collection of unrelated product logic.
That became a recorded design amendment on September 25. It sounds straightforward now, but it was an important step from “another frontend calling the same endpoint” to an architecture we could reuse responsibly.
Shared experience and execution, with product-owned tools and evidence.
We also talked about testing before implementation. I specifically asked what would happen to the existing BizOps AI regression tests. Those tests were evidence about an application we already depended on. They had to remain part of the refactor's acceptance, alongside the new Intersection tests.
The September 23–25 design ledger records these decisions and the testing additions. It also records a detail worth preserving: the initial independent reviewer was Opus 5. The later Opus 5.5 review was a separate, explicitly requested step.
September 26: I asked Claude to challenge the plan
This was a large enough refactor that I wanted a second opinion before Codex made extensive changes. I asked for a Claude-colab session using Claude Opus 5.5 Extra, with a clear division of responsibility: Codex would own implementation; Claude would independently review; fundamental product decisions would come back to me.
I wanted the implementation to work “one-shot.” Looking back, that was an ambition, not something a design review could guarantee. The review record correctly distinguished approval of a plan from proof that code, deployment and live answers worked.
The collaboration was structured. Codex sent bounded questions and patches to a persistent local Claude session. Claude inspected the code in read-only mode and returned findings. Codex checked those findings, made corrections and brought the revised work back. Claude's recorded source review did not become a claim that Claude had run tests; Codex remained responsible for producing execution evidence.
The working relationship: one implementation owner, an independent reviewer and explicit human product decisions.
Claude found real problems with assumptions in the plan. Financial context gates assumed a business unit. Some conversation paths discarded everything except message role and text. Existing chart generation did not yet provide the deterministic evidence boundary the plan described. Browser assets and independently deployed hosts introduced compatibility concerns.
Those findings mattered because a map area is not a business unit. We could not make geographic chat correct by giving it a fake BU and hoping the financial infrastructure would accept it. Nor could we preserve the meaning of “here” in a conversation if previous messages lost their original area context.
The revised plan reached approval in review turn five. I then reinforced what the first increment meant: get the chat interface and plumbing working, and learn from actual questions which additional tool functions would make the experience compelling. The direction was broad, grounded question answering. It was not a demand to anticipate every possible question before anyone used the feature.
The weekend: the difficult parts were below the chat window
Codex worked through the refactor in small slices. Much of that work was invisible from the map: conversation ownership, cancellation, run recovery, trusted application identity, operational attribution, storage writes and packaging.
One reviewer finding captures why some of this care was necessary. In an early cancellation change, Claude noticed that shutdown would immediately terminate running chats instead of allowing the existing grace period. The new test was effectively asserting the regression. If we had accepted it, a deploy or restart could have interrupted BizOps answers in a way the original application had not.
The correction needed a more precise invariant: a user-visible timeout is not proof that the underlying work has stopped. A request cannot release its capacity slot while its provider or database work continues unaccounted for. At the same time, a normal shutdown should preserve the established opportunity for an in-flight answer to complete.
That was one example among many. Codex also had to prove real-sized persistence writes, not just dry runs; installed deployment packages, not just imports that happened to work in the repository; and the existing application's behavior, not just the new geographic path.
The implementation ledger reaches Claude review turn 214. The final recorded review was still examining whether large OSM source collections remained reachable through bounded pages without losing their identities or provenance. This was considerably more than a second model saying the architecture looked reasonable.
It also became a very long review process. Early on September 28, in my local time, I told Codex:
“Continue solo without Claude.”
The ledger marks that change explicitly. Subsequent implementation and self-review belonged to Codex; it does not claim Claude approval for work performed after that boundary.
September 28–30: I had to bring the conversation back to delivery
After the weekend, I asked:
“What do you mean by ‘nothing was deployed’? You’ve worked the whole weekend. What has been achieved and what is still to complete?”
That was the tension in the task. There was substantial engineering progress, but I still did not have the live feature I was waiting for. Code completed, tests passed, packages verified and production available were separate milestones. Codex was careful about that distinction, but the distinction did not make the delay less frustrating.
I kept making product decisions as issues emerged. The AI had to generate tables in chat. Answers had to be grounded in the dataset. If it could not answer, it needed to say so honestly. I also asked Codex to relax the existing query deadline to 60 seconds so we could observe completed queries before tightening limits.
Boundary accuracy threatened to become another dependency for everything else. I separated that work so Codex could continue the AI delivery. The same principle applied more broadly: we needed to distinguish a missing product capability from a reason the shared engine could not be launched.
On September 30, I was more direct:
“I am waiting for you to complete now for 4 days on this deployment. I am the only user of Intersection Intelligence.”
My instruction was to be more willing to go live, provided the existing BizOps AI experience had not regressed. I accepted slow performance for the first Intersection release. We could improve responses and tune limits using actual interactions.
That was a specific product decision for a single-user first increment. Authentication, data integrity and BizOps compatibility still mattered. The change was in how we treated response-quality and performance findings: they no longer automatically meant rolling back the infrastructure we were trying to prove.
The conversation even contains a small reminder that this was work happening around ordinary life. I interrupted the flow for my commute home, then asked Codex to continue where it had left off. The task persisted through those interruptions; elapsed days were not uninterrupted autonomous working time.
September 30: the feature finally became something I could use
The launch record puts the first successful canonical release at 12:36 UTC on September 30, with independent verification immediately afterward.
The useful evidence was concrete. Signed-in native chat produced a table covering eight countries, accepted feedback and answered a follow-up about Botswana. The answers survived reloading the page. The conversation appeared in the existing BizOps Chat Logs. Agent Composer was absent from Intersection, as intended.
That established something we had been discussing for a week: a separate application could use the shared AI infrastructure and central operations while retaining its own context and tools.
I congratulated Codex, then changed the task from getting the engine into production to exercising it and fixing the issues we had deliberately allowed ourselves to investigate live. Shipping the first increment gave us a better source of priorities than another hypothetical catalogue of questions.
Sandton was a detour worth taking
Meanwhile, Claude had worked on the locality boundary problem in a separate branch. The issue was not simply bad geography. An ArcGIS GeoJSON export had represented some multi-ring features incorrectly, including holes that became filled polygons. The fix acquired Esri JSON, assembled rings according to orientation and checked remaining repairs against the source hierarchy.
Codex integrated that work and preserved its provenance in the AI projection. The release record reports 36,070 source-ring boundaries, 77 verified repairs and zero unavailable boundaries. Those were still Census 2011 boundaries. A geometrically valid repair could not establish that a polygon matched today's understanding of a suburb.
Once localities such as Sandton and Constantia existed in the dataset, we discussed how they should fit into the map. My concern was that a useful enrichment could distract from the application's purpose. This was controller and ownership intelligence, not an invitation to build a different kind of map product.
We brought locality into the existing navigation, with a feature flag and live scenarios. The map remained the starting point; the AI could use the selected place as context. Later conversations showed that locality support still had to be correct for each tool and record population. Being integrated into the map did not mean every query could inherit the selected locality indiscriminately.
October 1: I reminded Codex that reuse included the UX
After the infrastructure was live, I looked at the chat experience itself. BizOps already had the circular launcher, a panel that snapped smoothly into the page, and rich formatting for lists, tables and graphs. I had meant for those to be reusable too.
This exposed a gap between the plumbing we had focused on and the experience I had in mind. I did not want every future application to rediscover how to build the same chat panel.
Codex followed through by putting the shared launcher, docking, resizing, responsive shell and conversation components into the reusable client. Each host retained its page geometry and product adapters. For the native map, chat assets remained deferred until needed, and panel resizing avoided repeatedly triggering map work during the drag.
On October 2, we also added five contextual starter questions. They use the area, record or corridor already selected by the user, instead of making someone restate the page they are looking at. They do not query evidence simply because the user browses the map.
That was a more complete interpretation of Signal AI as a service: a familiar experience as well as a reusable engine.
Under the hood: one platform, multiple applications
Looking underneath that familiar chat panel shows what we actually made reusable. I wanted another application to inherit the experience and the operational infrastructure, while keeping control of its own intelligence. This detailed view, added on October 5, shows the boundaries we arrived at.
Open the diagram for the full-resolution image. Solid application paths describe the current architecture; dashed paths show how another application could join it.
At the top, each application hosts its own instance of the same client package. The launcher, panel, conversation controller and rich rendering are shared. The host supplies the selected business or geographic context, its authenticated transport and its page layout. Reusing the interface does not mix accounts, conversation scopes or permissions.
A submitted question passes through the application's authenticated bridge to Signal AI. The service resolves trusted application identity and canonical permissions, admits the work, manages the run and orchestrates the model's tool calls. Answers and evidence return through that authorized path. Durable run metadata, events and transcripts support recovery; provider activity, tool activity, feedback and health feed the existing platform operations experience. I do not need a second AI health console for the map.
The separation becomes most useful at the tool boundary. BizOps retains its existing financial tools and data paths. Intersection's adapter supplies its own read-only tools, evidence projections and authenticated query client. Those queries reuse the native map service, verified local query stores and reviewed evidence. The model does not receive direct database credentials, and the refactor does not create a second map database. The diagram also preserves the incremental nature of the work: we did not rewrite every existing BizOps path into the new adapter pattern, and Composer remains a BizOps feature.
For a future application, the work would be to define its context, permissions, tools and evidence, register a trusted adapter and host integration, and prove the shared lifecycle and operations still work. It would inherit the chat experience and engine. That is a concrete extension path, not a claim that a third application is already deployed—or that adding one requires no engineering.
The questions after launch were more useful than the questions before it
The Quebit conversation was particularly revealing. I asked which intersections had Quebit installed. The assistant searched the provider field, found zero matches and carefully explained the large unknown-provider population.
The qualifications were useful, but the search had still missed the relevant dimension. I pointed out that Quebit was recorded under adaptive technology. The later search found seven positioned installation records. Four links to junctions were candidate and three were ambiguous. That supported a claim about recorded installation evidence; it did not establish seven confirmed junction installations.
The experience gave us a concrete sequence of improvements: discover the right equipment fields, apply locality only to supported record populations, honor an explicit request to search more widely, preserve link qualifications and expose original source references.
Other questions led to more direct ownership lookups and better investigation continuity. We corrected “remaining records” versus “next page” wording. We made original inventory references distinguishable from rows in supporting documents. These were small changes compared with the engine refactor, but they were directly connected to whether a user could follow and trust an answer.
There is still a failure we should not hide. In the October 2 starter acceptance, the assistant found seven Quebit records and showed the right four-plus-three table, but its opening narrative said two records. The record retains that as an incorrect answer. A correct tool result or table does not validate every generated sentence around it.
We also had to stop treating every change like a data release
The long preparation wait after a restart prompted another discussion. Was Codex rebuilding South Africa whenever we changed an AI feature?
The distinction turned out to be important. The host was preparing a local copy of an existing database, not rebuilding the national geography. That still imposed a real wait, but it called for a different solution.
I asked for the chat launcher and panel to stay hidden until preparation completed. I also allowed a worst-case network-storage fallback after an operational copy failure, with a degraded-performance notice and the existing Operations reporting. Integrity failures remained closed.
Readiness states keep operational fallback distinct from normal local serving.
Later, when relaxing tool-call limits led to another build, I challenged whether these should be configuration settings. The repository still has different kinds of policy in different places; it would be inaccurate to claim every limit is now dynamically configurable. What we did establish was a release path that distinguishes browser-only, application and data changes. A UI-only change should not restart the native worker or copy the country stores.
How much work did this conversation consume?
For this retrospective, Codex inspected the local session records and produced a sanitized metrics snapshot, frozen at 14:00 UTC on October 3, before this rewrite. The measured thread includes the initial vision review, AI work, locality detours, validation, cleanup and documentation. It is not a clean measurement of feature implementation alone.
On narrow screens, scroll the table horizontally to read all columns.
| Recorded measure | Value | What it means |
|---|---|---|
| Conversation window | September 23–October 3 | More than ten elapsed days, including pauses, waits and interruptions. |
| Codex response usage records | 13,972 | Unique response IDs with usage telemetry; not 13,972 human messages. |
| Codex input tokens | 2,187,110,924 | Repeated model input across the thread, including cached context. |
| Cached input within that total | 2,147,430,784 — 98.19% | A subset of input, not extra tokens to add again. |
| Uncached input | 39,680,140 | Input less reported cached input. |
| Codex output tokens | 7,755,856 | Includes reported reasoning output; not all user-visible prose or code. |
| Total from response usage records | 2,194,866,780 | Input plus output, not unique text produced. |
| Context compaction records | 132 | Recorded compaction events in the parent conversation. |
| Main Claude review sequence | Reached turn 214 | The ledger includes design, implementation review and re-review. |
| Main Opus 5.5 session cost reported by Claude CLI | US$285.01 | Session telemetry, not a verified invoice or total project charge. |
The billion-token figure needs that explanation. It does not mean Codex wrote billions of tokens, or that billions of fresh tokens were billed at an uncached rate. The same long-running context was presented repeatedly, and the telemetry reports most input as cached.
There is an accounting caveat too. Summing unique token_usage_record entries exactly matches that record stream's final cumulative thread total. A separate event_msg/token_count counter reports 2,161,954,216 total tokens—32,912,564 fewer. The snapshot preserves both. We have not established why those counters differ, so these are explicitly local telemetry figures, not audited billing totals.
For Claude, Codex used the last cumulative cost state rather than adding repeated snapshots. That state reports approximately 897.83 million cache-read input tokens, 6.69 million cache-creation input tokens, 22,391 ordinary input tokens and 2.59 million output tokens for the main Opus 5.5 session. It excludes the earlier Opus 5 planning session and other sessions, including the separate boundary work.
We did not find a reliable dollar-cost record for Codex in this evidence. Applying an assumed public API price would turn an unknown into a plausible-looking number. Consequently, we do not have a defensible combined development cost. The Claude CLI figure is useful, but it is neither the whole project's cost nor proof of the amount charged under a subscription. Production chat usage and infrastructure costs are separate again.
These measurements describe the size of the process. They do not prove that this was the most efficient way to deliver the feature. The review effort, repeated context and time before launch are all things worth examining in future work.
October 3: finishing included the files and the documentation
A task this long leaves more behind than code. Our laptop had accumulated temporary datasets, build artifacts, test outputs and worktrees. I asked Codex to inspect and clean them, then to record a standing rule: managing its temporary files must be part of completing the work. I should not have to discover the problem when disk space runs out.
I also asked whether the documentation actually described the architecture we now had. Some of it did. Other current rules still directed all AI tools into the financial executor, and some contracts still said the feature had not been deployed.
Codex reconciled 29 documents, added the architecture guide and diagrams, and linked current guidance to the release evidence. Historical records were retained as history. That matters for the next agent or developer: stale instructions can quietly undo the separation we spent days establishing.
What I take from our longest task
I now have a usable first AI increment inside Intersection Intelligence, backed by a reusable engine and chat experience. BizOps remains the operational platform. Intersection owns its evidence tools and geographic context. Future applications have a concrete integration pattern to follow.
I also have a much clearer view of my role when working with autonomous coding agents. I needed to define the product boundaries, insist on preserving the existing application, ask for independent review, answer unresolved product questions and challenge the point at which caution was delaying the first useful release.
Codex did the sustained implementation and verification work. Claude challenged assumptions and caught regressions. Neither could decide for me how much slowness was acceptable for a first release, whether a locality detour served the product, or when we had enough infrastructure to start learning from use.
The most useful questions eventually came from the live map: who owns this intersection, what equipment is recorded here, where did that fact come from, and what have we not established? Those questions are now driving smaller, more concrete improvements. That is where I want our next phase of work to stay.




No comments:
Post a Comment