The end of my benchmark tests of AI coding model's ability to represent a real world hourglass timer, in code, without using large gaming / 3d engines. Opus5.5 built this in one shot from a simple prompt:
Build me a single page application that is a digital twin of an hour glass timer. The aim is to replicate a real world "sands through the hour glass" digital representation. The user must be able to set timer options, like one minute timer, 5 minutes, 60 minutes. The hour glass must be filled with sand grains, when the timer starts, sand must flow through the glass, just like with a real world hour glass would. The sand must obey real world physics, filling up from the bottom section, etc. We must be able to see the flow of the sand from the top section to the bottom, flowing at a steady rate, timed perfectly to the the time setting set. Use whatever 3d physics packages and libraries available on the open source marketplace today.Code's on Github: https://github.com/khanmjk/HourGlass_Opus55
App, ready for anyone to use: https://khanmjk.github.io/HourGlass_Opus55/
For the full history of my benchmarking journey, check my blog: https://khanmjk-outlet.blogspot.com/search/label/HourGlass
This post marks the end of my benchmarking test. The models have come quite far indeed, however, Claude's Opus 5.5 is by far the best implementation. My test exposes model's ability for general knowledge, science, physics and simulation - and software coding expertise. Whilst some might argue my test is very simple, that models have been shown to create far more complex applications, gaming worlds and much more complex simulations -- true, I'm not debating this -- but -- my tests expose how far the gaps remain, still today - where model providers are releasing state-of-the-art, most advanced versions to date - scoring the highest in benchmarking metrics -- yet oddly struggle to start from a simple prompt and implement an hourglass digital twin that is usable. Only Claude Opus 4.5+ succeeded, with models from OpenAI and Google, failing miserably. As advanced as GPT 6 Astra is (and I use it every day for building enterprise apps), I'm amazed by how it failed my hourglass test. The point is that these models are not demonstrating consistency. Consistency earns trust. When trust is earned, adoption and engagement accelerates. What we need is a generally consistent model that is competent at a spectrum of tasks, not forcing the users to decide which model to farm the task to. I wonder how routers like OpenRouter would take my prompt and decide which model is best placed to build the Hourglass sim - maybe that's the next phase of my benchmarking test. Can we trust these routers enough on their decision-making abilities?
#AI #opus55 #hourglass #digitaltwin
No comments:
Post a Comment