This is a follow-up to I asked AI to count my carbs 27,000 times. It couldn’t give me the same answer twice. The full preprint is on Zenodo at doi.org/10.5281/zenodo.22879139.
When OpenAI launched GPT-6 Astra this month, its president Greg Brockman said it’s not unreasonable to feel we’re now in the AGI era, and then said he’d leave it to us to decide whether that qualifies.
Anthropic hasn’t used the expression AGI for Claude Fable 5.1, instead calling it the world’s most advanced model for coding and knowledge work, so it seemed only fair to put both of them through the same test I gave their predecessors back in April and see how human they really are.
But here’s the sneak preview. I’ve already decided. At counting carbs, it makes plenty of mistakes and in the wrong direction. Like humans, but not really like humans. So if being a bit like humans by making mistakes qualifies as AGI, then well done Anthropic and OpenAI, you’ve made it. But I’m not sure that’s quite what they meant…
It’s worth knowing what human looks like first. In one study of 50 adults with type 1, most of whom had been counting carbs for more than 20 years, people were off by about 15 g a meal, or roughly a fifth of what was on the plate, and most of those errors were undercounts. Another found that people estimating hospital meals were out by 28 g on average. On the five meals in my test where I know the answer with some confidence, Fable 5.1 was out by 7 g on average and Astra by 13 g, which is about a sixth and a third of the meal. My meals were smaller than the ones in that first study, and it wasn’t the same test, so it’s not a straight fight, but by that yardstick Fable 5.1 counts carbs about as well as we do, while Astra still has some way to go.
Having said that, human-level is a pretty poor bar for something that’s going to set an insulin dose, and as we’ll see, the models don’t get it wrong in the same way we do. We tend to undercount. Astra overcounted, and asking it the same question twenty times didn’t make the error go away, it just gave you the same wrong answer more reliably.
How we got here
This all started at ATTD in March, where I had 10 people send the same 6 photos to ChatGPT 5.2 to see whether it would give the same answers for the same pictures when they came from different people. It did a great job of failing, with portion sizes, doses and even what it thought the food was varying depending on who sent it.
To follow that up, in April I sent 13 food photos to four AI models around 500 times each and asked them to count the carbs. The headline was pretty simple: ask the same model the same question about the same photo, and you get a different answer. I finished that piece by suggesting that if an app sent each photo three to five times and showed you the median, a lot of that wobble should go away, and this article is really about testing that idea.
Things have moved on since April, though. Both Fable 5.1 and Astra have arrived without the option to set the temperature, which is the setting I’d used to keep the April models as close to deterministic as possible (Fable 5.1 throws an error if you try, and OpenAI’s migration guide tells you to take it out), and neither of them can run with reasoning switched off. That means any app built on them gets whatever sampling the provider chooses. So this time round I sent the same 13 photos, with the same prompt, 50 times each to Fable 5.1 and Astra at the lowest reasoning effort they allow, and to two of the April models, Sonnet 4.6 and GPT-5.4, with the temperature setting removed and nothing else changed.1
It’s worth being clear about how much evidence that really is, as the number of API calls makes it look a lot bigger than it is. Of the 13 meals, only five have a solid reference, two from packet labels and three where I weighed the food, and even those are only good to within a few grams. Three more were portioned but not weighed, the churros are my visual estimate from a restaurant, and four have no reference at all and are only there to test consistency. So everything I say about accuracy rests on five meals, and thousands of answers about five meals are still only five meals.
The chart below shows most of the story before we get anywhere near the statistics. GPT-5.4 put the cheese sandwich around 33 g too high and Astra was well over on the burrito, Fable 5.1 and Sonnet 4.6 both had the sandwich about 12 g low, and on the chilli, every model was high.

Asking more than once
The first thing to look at was whether removing the temperature control made the answers wobble more, and unsurprisingly, it did. Sonnet 4.6’s median variation went from 2.3% to 5.7% and GPT-5.4’s from 7.8% to 10.1%, with both more variable on almost every photo, while the two new models landed close to Sonnet 4.6, at 6.3% for Fable 5.1 and 5.7% for Astra. There’s a slightly awkward footnote to this. The April GPT-5.4 data had been collected in lots of separate batches, and the spread varied between them from 4.3% to 9.1%, which suggests that how consistent a model looks in one session may understate what you’d see across several.

Next came the averaging. I didn’t make any new calls for this, but instead simulated it by drawing batches of between 1 and 20 answers from the 50 that each model gave for each photo, and taking the median. For the stray answers, it works well. With a single call, somewhere between 3.4% and 17.5% of answers landed more than 10 g away from what that model usually said, and with the median of 20 that fell to 1.5% or less for every model, with five calls getting most of the way there for Fable 5.1 and Sonnet 4.6. The really big misses, the ones that would have meant more than 5 U too much insulin, disappeared too, at least in this data. It does need to be the median though, as with the mean one wild answer drags everything with it. It’s also worth saying that this assumes repeated calls behave like draws from the same pot, which the April batches suggest isn’t always true, so the real test would be fresh calls spread over several days.
So far, so good. The problem comes when you compare the answers with the reference values rather than with the model itself.
To keep this simple, when I talk about an overdose here, I mean that bolusing for the estimate at 1 U per 10 g would have given you more than 2 U of meal insulin too much. Nobody was actually dosed, and it ignores IOB and anything your loop would do next. On the five solid meals, Astra’s estimates did this 19.2% of the time with one call and 20.0% of the time with the median of twenty, while GPT-5.4 went from 34.7% to 40.0%. Those are averages over five meals, and one burrito can move them a long way, but the direction is clear enough. Averaging didn’t bring them down, and for GPT-5.4 it made things worse.

The reason is fairly obvious once you think about it. Averaging pulls the answer towards whatever the model usually says for that photo. If that usual answer is close to the truth, averaging cleans up the stray bad answers, but if it’s wrong, averaging makes sure you get the wrong answer every time. Astra’s typical answer for the breakfast burrito was 37 g too high, and its overdose rate went from 94.0% on a single call to 100% with twenty. Sonnet 4.6 on the chilli, 21 g high, went from 68.4% to 94.2%. On the other hand, GPT-5.4 on the stuffed pork loin was usually only 4 g high, and averaging took its overdose rate from 34.5% down to 5.0%. It works the other way round too, with GPT-5.4 undercounting the churros by more than 20 g in 73.9% of single calls and 99.3% of twenty-call medians, although that’s the reference I estimated by eye.
This is what worries me. With one call, a biased model gives you a number that jumps about a bit, and you might notice that something’s off. With twenty calls you get nearly the same number every time, and it’s very easy to take that consistency as a sign that it’s right.

Undercounting follows the same pattern on a smaller scale. Leaving the churros aside, no model was ever more than 20 g under on any meal with a reference, but at 10 g it’s a different picture, with Fable 5.1 low in about a fifth of its answers on the solid meals and Sonnet 4.6 in about two fifths, most of that down to the cheese sandwich. Averaging locked those in just as it did the overcounts. Too little insulin usually means a high rather than a low, which is less immediate, but it’s not nothing, particularly if you’re pre-bolusing precisely.
All of this also depends on your carb ratio. I used 1 U per 10 g as an illustration, but at 1 U per 5 g the same 10 g error becomes a 2 U overdose, and at that ratio every model overdosed in at least 15% of single estimates, including Fable 5.1 and Sonnet 4.6, which had been at zero at 1:10. At 1 U per 20 g, only Astra still had a meaningful rate. The fewer grams each unit covers for you, the more the same carb error matters. And what about us humans? As I mentioned earlier, we’re far more likely to underestimate, which is much less likely to put us into a hypo three hours after eating.
So which one is best?
On the five solid meals, Fable 5.1 had a mean absolute error of 7.0 g with almost no bias and Sonnet 4.6 was at 8.8 g, while Astra and GPT-5.4 came in at 13.2 g and 14.7 g, both overestimating by about 9 to 10 g on average. As a share of the meal, that puts Fable 5.1 at 17.5%, against about 21% for the adults in the Brazeau study.
It would be easy to treat that as a ranking, but it isn’t one I can stand behind. When I compare the models meal by meal, every difference has a confidence interval that crosses zero, and as a rough guide, confirming the 1.8 g gap between Fable 5.1 and Sonnet 4.6 would need something like 94 meals. I have five. Anyone claiming one of these models is clearly the best at carb counting on a test set this size is overreaching, and that includes this article. For what it’s worth, removing the temperature control didn’t make the April models any less accurate either, just less consistent.
The new models are better at knowing what they’re looking at, at least. In April, Sonnet 4.6 called the Bakewell tart a Linzer torte every time, whereas both Fable 5.1 and Astra got it right, along with the crema catalana. They have their own quirks though, with Fable 5.1 deciding the stuffed pork loin was chicken in almost half its answers, and Astra including the burrata salad from the next plate in the pizza photo. On the whole, neither made much difference to the carb numbers.
Some of the models also occasionally chucked back answers that simply weren’t usable. Astra, for example, wrote “forty” rather than 40 for most of its answers on the churros, which breaks any app that’s expecting a number. If you’re building one of these apps, it needs to check what comes back and tell the user when it hasn’t got an answer, rather than trying to guess.
Cost is probably the biggest reason that apps aren’t doing any of this. At September 2026 prices, and assuming you want the answer while you’re sat in front of the meal rather than in an overnight batch, a single call per meal to GPT-5.4, the cheapest model here, works out at about $2 per user per month for three meals a day. Asking five times, which is where most of the benefit of averaging came, costs about $16 a month on Sonnet 4.6 and $55 to $57 a month on either of the new flagship models. That’s before the app has paid for anything else, and it’s a lot to fold into a subscription, which goes a long way to explaining why apps aren’t doing it.
So, is it AGI?
If the test is doing a job as well as a person, then on carbs Fable 5.1 is near enough, and Astra, the one whose maker raised the question, is still learning. A purpose-built phone app got within about a gram of Astra’s average back in 2016, but let’s not spoil the party.
But you already had a person doing this job. You. You get it wrong too, but you know when you’re guessing. You know the chilli was a big bowl and the sandwich was thick-cut. The model gives you a number with no sign of whether it’s guessing, and if you ask it twenty times, it gives you that number twenty times.
Averaging is still worth doing, as it gets rid of the random bad answer, and in this data it got rid of the really big ones too. What it can’t do is anything about a model that consistently gets a meal wrong. So if you’re using one of these tools, it’s worth looking at the number every time and asking whether it makes sense for that plate, and being especially careful if each unit only covers a few grams for you. If an app always gives you the same answer for a meal and it always feels a bit high, that’s bias rather than scatter, and asking again won’t fix it. If you’re building one, use the median and tell the user when you haven’t got a usable answer, and given how differently the models behave, it’s worth restricting it to the ones you’ve actually tested.
On these 13 meals, the errors were big enough that I wouldn’t dose from either new model without checking. Nor the old ones.
1 I used Claude Code to help build the benchmark and the analysis. As Anthropic also makes two of the models tested, the prompt and all of the code that reads the answers are identical for every model, and everything is public on GitHub. ↩
Preprint: Street T (2026). Carbohydrate estimates from food photographs by Claude Fable 5.1 and GPT-6 Astra: a September 2026 update to a reproducibility benchmark. https://doi.org/10.5281/zenodo.22879139
Code and data: https://github.com/tim2000s/llm-food-benchmark-2026-09
Original April article: I asked AI to count my carbs 27,000 times
Other sources
Brockman on the AGI era: Gizmodo, OpenAI claims we’re in the ‘AGI era’ with release of GPT-6 Astra; TechTimes, GPT-6 Astra goes live (September 2026).
Anthropic (2026). Introducing Claude Fable 5.1 and Claude Mythos 5.1.
Brazeau AS et al. (2013). Carbohydrate counting accuracy and blood glucose variability in adults with type 1 diabetes. Diabetes Research and Clinical Practice. pubmed.ncbi.nlm.nih.gov/23146371
Rhyner D et al. (2016). Carbohydrate estimation by a mobile phone-based system versus self-estimations of individuals with type 1 diabetes mellitus: a comparative study. Journal of Medical Internet Research. pmc.ncbi.nlm.nih.gov/articles/PMC4880742
Discover more from Diabettech - Diabetes and Technology
Subscribe to get the latest posts sent to your email.
Leave a Reply