People with type 1 diabetes have been asking chatbots to look at their data for a while now, and this year I have spent a good deal of time finding out what happens when you ask the same question more than once. In April I asked four models to count the carbohydrates in thirteen photographs, 26,904 times in all, and could not get the same answer twice. A week later I gave five models a week of raw pump data from three users and asked for settings, and the settings came from the textbook rather than the data. In both cases the person doing the asking was on their own. You exported a file or took a photograph, you pasted it in, you read what came back with the scepticism it deserved, and if you were sensible you asked again to see whether the answer held.
Someone has now built the product. TuneMyBG is an Android app, £4.49 a month, that connects to your Nightscout instance, packages your last 7, 14 or 30 days together with your AndroidAPS profile into a structured file, and hands you four prompts. You open a fresh conversation in whichever chatbot you like, attach the file with the first prompt, send the other three in turn, and paste the final answer back. The app checks it, sends you back to the chatbot with a correction prompt if the check fails, and otherwise shows you a primary recommendation, a keep or change verdict on each of 23 settings, a meal strategy and a list of actions whose progress you can track. The listing calls it “AI ready Nightscout data for safer AAPS tuning”. It does not call itself a digital endocrinologist. I am calling it that, because a service that reads your data and tells you which of your insulin settings to change is what an endocrinologist does, and this one charges less than a coffee a month and never has a waiting list.
The pitch is that structure makes the chatbot safer. The data arrives tidy, the prompts are fixed, the answer has to fit a schema, and the app checks it before you see it. That is a reasonable idea and I wanted to know whether it works, so I did what the app asks a user to do, fifty times over, against eleven different models, on my own data.
Fifty conversations, one record
The method was simple. Take the app’s own package and its own four prompts, start a fresh conversation, run the sequence, and judge the answer the way the app judges it. I had to work out how the app judges answers by pasting 25 real chatbot outputs into it and watching what it did, because the source is not published; the checker I built from that agrees with the app on all 25. Then repeat, fifty times per model where the cost allowed, with Gemini 3.6 Flash, Gemini 3.1 Pro, Gemini 3.5 Flash-Lite, GPT-5.6, GPT-5.4-mini, Claude Opus 5, Claude Sonnet 5, Claude Haiku 4.5, DeepSeek V4 Pro, Grok 4.6 and Llama 4 Maverick. The app recommends no model, so any of these is a legitimate choice for a paying user. My record was fourteen days, time in range 86 per cent, mean glucose 121.6 mg/dl, 5 per cent below 70, and a cluster of lows in the hour after midnight that the rest of the night resolves on its own, which on a ten hour duration of insulin action looks a lot more like the tail of evening insulin than like a wrong basal rate.
The first result is the reassuring one. The models read the record well. Between 96.8 and 100 per cent of the figures they quoted were actually in the file, and the three models that went after my safety limits all did the same sum, a 25 unit maximum insulin on board against roughly 20 units a day. Whatever else is going on, the structured package does most of its job: eight of the eleven models argued nearly every change they made from a figure in my record, Gemini 3.1 Pro two thirds of its changes, and no model quoted a number that was not there. Llama was the exception, changing settings with no figure at all in 53 of its 56 changes.
There is a catch in that sum, though, and it is the app’s. Twenty units a day is not my insulin use; it is my boluses. The package totals only the bolus and SMB records, carries no basal delivery at all, and my scheduled basal adds another 16 units a day, so my real total is nearer 36. The models split on noticing. Gemini 3.6 Flash took the bolus-only figure as my daily total in 49 of 50 conversations and cut my safety limits in 48 of them partly on the strength of it; Opus added the basal back and reached about 36 in all eleven of its accepted conversations; Haiku flagged the missing basal in 19 of 20. The tidy package feeds every model the same understatement, and whether your reviewer notices depends, once again, on which one you opened.
The models that caught the app’s understatement caught it with knowledge the workflow tells them to set aside: that a closed loop delivers basal insulin, and that nobody runs on twenty units a day with my profile. The textbook knowledge the settings study blamed for anchoring turns out to be the only thing a model can check your data against, and a design built on tidy data in and obedient reading out selects for the models that never run that check.
The insulin total is the biggest trap in the file, and it has company. The package gives my glucose by hour twice over, once in UTC and once in local time, and thirty conversations placed my night lows at the wrong hour of the clock. The only mention that I use diluted insulin is a free-text profile name, which ninety conversations spotted and the rest never saw. The loop’s own reasoning strings are in mmol/L inside a file that is otherwise mg/dl, a trap two conversations quoted and none acted on. Each split is the same story: the app decides what the model can see, and the model decides, one draw at a time, how carefully to look. Even DeepSeek’s sensitivity factors from 30 to 180 look less like invention once you notice that the package’s own dynamic sensitivity values run from 38 to 124; it was arguing from my record even when it looked like it made no sense, and couldn’t distinguish between profile and dynamicISF.
The second result is the one you are paying for, and it is not reassuring at all. Ask Gemini 3.6 Flash fifty times and it lowers the maximum IOB limit in 49 of them, which sounds consistent until you look at the number it proposes. Seven units in 27 conversations, six in some, four in others, fifteen in one, and the maximum basal alongside it anywhere from 2.5 to 4 units an hour. That is eighteen different pairs of safety limits from one model on one record. Gemini 3.1 Pro produced seventeen pairs and Opus five, forty distinct answers between them to a question that has one right answer or none. Four other models, GPT-5.6, GPT-5.4-mini, Sonnet and Grok, looked at the same file and left the limits alone in every conversation. DeepSeek changed my sensitivity factor in 20 of 48 conversations, to fourteen different values between 30 and 180 mg/dl per unit against my current 56, so that one run would have doubled the strength of every correction and another would have cut it by two thirds. Llama raised my midnight basal in sixteen conversations and lowered it in seven.
The headline recommendation flips too. Four times out of five Gemini 3.6 Flash told me to raise my overnight targets and led with “target”; the fifth time it left the targets alone, cut the midnight basal and led with “basal”. Both versions cited the same figure, 18.5 per cent of readings below 70 at one in the morning. Same evidence, two stories, and the one you get depends on which way the dice fell when you pressed send.
The overnight lows are where the difference bites. Opus and Sonnet worked out, in all but one conversation between them, that the lows were the tail of earlier insulin, and left the night-time settings alone. Gemini 3.1 Pro worked out the same thing in 40 of 50 conversations, named the insulin stacking, and then raised the target anyway as a cushion. Llama changed the basal in half its conversations and in 18 of them gave no reason at all beyond “No strong evidence to change”. Only one of those three readings acts on what the record shows, and all three arrive on your screen in the same format with the same confidence.
The image below shows the variation on some models and not others, and given the sparsity of change in everything except Gemini-based systems, raises the question as to how the app was a) built; and b) tested.

Which universe are you in?
Put yourself in the position of a paying user. You open Gemini, run the four prompts once and get an answer: lower the max IOB to 7, the max basal to 3.5, raise the overnight targets. It arrives in the same format and with the same confidence as every other answer, and the app stamps it as checked. What you cannot see is that you have been handed one cell from the barcode above. Had you pressed send a minute later you might be in the universe where the limit is 4, or 15, or the one where the targets stay put and the basal is cut instead, and from inside any one of them the others are invisible. Open a different model and the universes multiply: one where nothing changes, one where your sensitivity factor triples, one where the midnight basal goes up and one where it comes down. Each is internally consistent, each cites your numbers correctly, and each would pass the app’s check. The moment you pick a model and press send you step into one of them at random, and nothing on the screen tells you which.
This is how these systems work, whatever the product and whichever the model. A language model produces text by sampling, and the chat interfaces give you no control over it, so the same question asked twice is two draws. Where the evidence points one way the draws agree, which is why the models all read my record correctly. Where the evidence admits more than one story, as it usually does with a fortnight of pump data, the draws disagree, and the app outsources your settings review to a single roll of the dice inside whichever model you happened to open, then hides the dice.
And the universes differ in more than their numbers. In some of them your reviewer audited the file before advising you: it noticed that the app’s insulin total was missing your basal, rebuilt the real figure, and judged your safety limits against that. In others it took the understated total at face value and cut your limits accordingly. Same app, same data, same prompts; whether anyone checked the arithmetic depended on which model you opened, and sometimes on which of its answers you happened to draw. So the one decision the app leaves entirely in your hands, which chatbot to use, turns out to be the decision that matters most, and it is the one the app gives you no help with at all.
What the check actually checks
None of this was caught, because the app’s check does not look at content. It looks at whether the 23 rows are all present, whether the current values were copied correctly, and whether the headline focus is one of five permitted words. A sensitivity factor of 180 passes as easily as 56. A meal strategy step marked “completed” on a record that contains not a single carbohydrate entry (I do not log them) passes, and three models marked it completed in all or all but one of their conversations. The prompts contain rules about confidence words and required sections, but they are instructions to the chatbot, not checks, and the cheaper models broke them in a quarter to four fifths of the answers the app accepted.
The check does have teeth in one direction. When a model produces something that is not valid JSON, which Opus did in nine conversations out of twenty, usually by dropping one closing brace, the app says “Paste a valid JSON response from AI” and offers nothing further. Haiku was still being rejected after the correction prompt in 27 of 50 conversations, for repeating a phrase in the focus field that the app will not accept. So the app is fussy about shape and indifferent to substance, which is precisely the wrong way round for something that is going to be read as a settings review.
A medical device, on its own description
The disclaimer is meant to cover what follows. The onboarding screen says the app does not replace medical advice, that you should not change settings solely on an app result, and that you use it at your own risk. That is a contract between you and the developer, and the regulation does not ask about contracts.
The EU Medical Device Regulation defines a device by what the manufacturer intends it for, and reads that intention from the label, the instructions and the marketing. “Safer AAPS tuning” and “possible profile improvements”, with a tracker for whether you applied them, is a therapeutic purpose. The Commission’s guidance on software (MDCG 2019-11) says in terms that software which processes or analyses medical information for a medical purpose qualifies, that it does not matter where the processing happens, and it lists “software that provides insulin dose recommendations to a patient regardless of the method of delivery” among its examples. The argument that the app only organises data and the chatbot does the thinking does not survive that: the prompts, the schema, the check and the screen are the developer’s, and they are what turn a chat into a settings review. The argument that the user decides does not survive it either, because Rule 11, the software rule, exists for information that a person acts on. Rule 11 puts software used for therapeutic decisions in Class IIa, and in Class IIb where a wrong decision can cause serious deterioration. Severe hypoglycaemia and ketoacidosis qualify. Either class needs a notified body, a quality system, clinical evidence for the claims and post-market surveillance. Neither can be self-declared, and the Play Store listing carries no medical device label. I have set the full reasoning out in a commentary submitted alongside the study, and the classification is ultimately for the developer and the regulator, but I cannot find the reading of the texts under which this is not a device.
Then there is where your data goes. The package is your health record for a fortnight, identifiable by its timestamps and your profile, and the workflow has you paste it into a consumer chat service whose terms may let it train on the conversation unless you have opted out. You make that transfer, not the app, and the onboarding says nothing about it.
What a responsible version would look like
People ask models to look at their data, I do it myself, and the structured package genuinely helps the model read it. The objection is to selling the answer without owning it, and owning it means doing the things a device manufacturer does. Pick a model and pin the version, so the product is the same product next month. Test it, so you can say what it does with a record like mine. Run the analysis several times and show the spread, so the user can see that “lower max IOB to 7” is one draw from a range that runs from 4 to 15.
Check the content of the answer, not just its shape. Say on the listing what the app is and what class it is in. Say where the data goes. Make sure that what you’re giving to the LLM is what you think you’re giving to the LLM. Insulin? Both bolus and basal… If you’re leaving the model choice up to the user, make very sure that you know how it will respond across any model they will choose, not just the ones you’ve tested against.
And if you’re going to vibe code an app, be very sure you really know what it is doing.
The study found what a manufacturer who did those things would have found first, and the developer could have done it for what I spent, which was about £150 in API credit.
Until then, the advice that DTN published in January still stands: any suggested change to insulin settings from a general purpose model needs a careful sense check before you act on it. This app makes that harder, because it dresses the chatbot’s answer in the form of a review and stamps it as checked. You are left to check it against your own judgement, which is what you were paying to supplement.
An endocrinologist who gave you a different set of numbers every time you asked is not one you would keep paying. Neither is this. Keep the data, keep the money, and if you want a second opinion from a chatbot, ask it three times and see whether it agrees with itself.
The study and its companion commentary have been submitted for peer review and are posted as preprints here and here and on SSRN; the harness, every transcript and the anonymised package are in the public repository. The app’s developer has no connection with me and I paid for the subscription.
Writing these articles and firing off multiple questions to LLM APIs doesn’t come for free. If you’d like to support more of this, please consider popping a donation into Stripe to support more of these investigations.
Discover more from Diabettech - Diabetes and Technology
Subscribe to get the latest posts sent to your email.
Leave a Reply