Can an AI Health Coach Be Trusted With Your Health? A Clinician’s Honest Take

follow jackie

Board-certified Family Nurse Practitioner and Nationally Certified Menopause Practitioner with 15+ years in clinical practice, specializing in perimenopause, menopause, hormone health, and sexual medicine. 

Hi, I'm Jackie Giannelli, FNP-BC, NCMP

Medically reviewed by Jackie Giannelli, FNP-BC, NCMP.

We have to talk about a real, growing issue right now: using AI as a health coach. I know you’ve already asked it something. Maybe about the hair in the shower drain, or the anxiety that arrives at 4 pm for no reason. A little vaginal discomfort. It answered in four tidy, super confident paragraphs. So here is the real question about using an AI health coach: how would you know if it was wrong?

I use AI every day. And I also read the papers on where it fails, and they are not flattering. Both things are true at once, and the reality lands somewhere in the middle.

Before you go further: if you have bloodwork sitting in a portal you have never really read, my free Lab Guide walks you through the markers that actually matter for women 35+ and what each one is telling you. Get the Lab Guide → It pairs well with everything below, because AI is far more useful when you feed it your own numbers than when you ask it questions in the abstract.

What an “AI health coach” actually is right now

There are generally three different types of AI health coaches:

  1. A general chatbot
  2. AI health products, the ones with the app and the subscription and the onboarding quiz
  3. Using a general model deliberately, with a structure around it, on data you have already gathered. This is what I teach you to do in the Longevity Lab. Click here to get on the waitlist, we’re starting in early October.

All three are the same underlying technology, which is a large language model (LLM). LLMs predicts the next likely piece of text based on patterns in what it was trained on. It is very good at that. What it is not doing, at any point, is reasoning about your body. It has no idea what your ferritin was, whether you have a family history, or that you have been feeling like this for three years. It simply knows what an answer to a question like yours usually looks like.

The AI safety research

In February 2026, a group of physicians published an evaluation of four leading chatbots against 222 real patient-posed medical questions, across primary care, women’s health, and pediatrics. Eight hundred and eighty-eight answers, graded by doctors. (1)

The results: problematic responses ranged from 21.6% to 43.2% depending on the model. Outright unsafe responses ranged from 5% to 13%. (1)

Sit with the low end of that for a second. Five percent, from the best-performing model. The authors ran the arithmetic themselves: if only 5% of the roughly 43 million medical questions asked monthly are answered unsafely, that is more than two million unsafe responses. (1)

The failure modes were not exotic either. Missing the history-taking that any clinician would do first. False reassurance given without knowing the context that would change the answer. Advice that was actively dangerous for specific people. Their conclusion was that current safeguards are not sufficient. (1)

The truth remains that there’s a serious desire to get health answers from a chatbot, and the reasons for that are a discussion about the medical system, access, and insurance that we’ll have to talk about another day. About one in four American adults used an AI tool for health information in the last thirty days. (2) So “don’t use AI for your health” isn’t a realistic ask.

So let’s talk about how to use it a little bit better, shall we?

Why the issue is more pressing for women

The research it learned from left you out

AI learned medicine from the published literature, and the published literature has a very specific hole in it.

Only around 5% of clinical trials report sex-disaggregated data. (3) Not 5% of trials that exclude women, 5% that bother to report results separately for the women who were in them. Everything else gets averaged into a single number that describes a body that may or may not resemble yours.

Then there is the funding. Of private healthcare capital globally, roughly 6% goes to conditions affecting women, and 90% of that concentrates in cancers, reproductive health, and maternity. Conditions like endometriosis, menopause, and PMOS together receive under 2% of private healthcare funding. (4)

And the downstream effect of all of it: across a population-wide analysis spanning two decades, women were diagnosed later than men for more than 700 diseases, by an average of four years. (3)

None of that is the AI’s fault. But it’s the reality in which AI was created and trained on.

What that looks like in an answer

Ask a chatbot about a fluctuating hormone, and you will get an answer built on the strongest available evidence, which for midlife women is thinner and older than you would expect. Ask it about a symptom cluster that shows up in perimenopause, and it will frequently route you toward thyroid, or iron, or stress, because those are the well-studied explanations with the most text behind them.

Sometimes it is right. Thyroid and iron genuinely belong on that list, and I check both. But the pattern I see is that a generic AI health coach reaches for the well-documented answer over the one that fits the woman, because the well-documented answer is what it has the most of.

You are the person for whom “the average patient in the literature” is least likely to be you.

Where an AI health coach earns its place

Everything above is about asking AI to be a doctor. It is bad at that. It is genuinely, usefully good at four other jobs.

Organizing what you already have

You have labs across three portals, a symptom pattern you have described out loud maybe forty times, an app with two years of sleep data, and a family history you can recite. They have never all been in one place, because putting them there is four hours of tedious work.

That is a text-organizing problem. It is exactly what this technology is for.

Translating the language

A lab report is written for the clinician who ordered it, not for you. Asking an AI health coach what a marker measures, what raises and lowers it, and what a value in your range typically indicates is a translation task, and translation is the thing this technology does best.

It is also low-risk, because you can check the answer. Ferritin either is a measure of stored iron, or it isn’t. That is a fact you can verify in thirty seconds, unlike “what should I do about my ferritin,” which depends on you.

The rule of thumb I use: if the question has one right answer that exists somewhere in a textbook, ask away. If the answer changes depending on who is asking, that is a question for a person who knows you.

Building the question list

This one changes appointments more than anything else on the list.

Twelve minutes is not enough time to work out what to ask while you are also being examined, remembering a symptom you meant to mention, and doing arithmetic about dates. So you leave having covered the thing that was easiest to say out loud, which is rarely the thing you came in about.

Walking in with six specific written questions, ordered by priority, is worth more than any single test you could request. Ask AI to draft that list from your own information. Then cut it down yourself, because you know which three actually matter and it does not.

A version I like: ask it what a clinician would most want to know about your situation that you have not written down yet. The gaps it finds are usually pretty good.

Seeing your own patterns

Give it your last three years of cycle dates and your symptom notes and ask what moves together. Not “what is wrong with me.” It will answer that question poorly. “What co-occurs in this data” is a pattern question, on your data, and it is the one worth asking.

I watched AI read a patient’s biomarkers and pick up a connection three specialists had not made. That happened, and I have never stopped thinking about it. But, it also happened with a clinician in the room deciding what to do about it.

Five rules I’d give you before you ask it another thing

  1. Give it your data, not just your symptoms. A question with your actual numbers attached produces a fundamentally different answer than the same question asked in the abstract. If you change one thing about how you use it, change this one.
  2. Never let it be the only voice. Its output is a draft for a conversation with a licensed human, not a substitute for one. Bring it in to your next appointment. I would rather see what you asked than not know.
  3. Make it show its work. Ask what it is basing an answer on and how confident it is. When it cannot tell you, that is information.
  4. Ask what it does not know about you. “What would you need to know about me to answer this properly?” is the best prompt in health, because it forces the history-taking step the research found missing. (1)
  5. Treat confidence as a style, not a signal. It writes with the same certainty when it is right and when it is inventing. The tone tells you nothing. Verify anything you would act on.

The version of this I actually recommend

I am a nurse practitioner. I am not a coach, and I am not going to tell you that a chatbot is going to sort out your hormones, because it isn’t.

What I do believe, and I stand by this POV: you already have most of the information you need. It is scattered across portals and apps and your own memory, and reading the connections across it is a separate skill from collecting it. That is not a data problem. It is an interpretation problem.

An AI health coach, used with structure and with a clinician’s framework around it, is very good at interpretation. Used as an oracle, on questions about a body the research barely studied, it is exactly as risky as the numbers above suggest.

The difference between those two outcomes is not the technology. It is what you feed it and what you do with what comes back.

Frequently asked questions

Is an AI health coach safe to use?

It depends entirely on what you use it for. A 2026 physician-graded evaluation of four leading chatbots found unsafe answers in 5% to 13% of responses to real patient questions, depending on the model. (1) That is a meaningful risk for diagnostic or treatment questions and a very low risk for organizing your own records, translating a lab report, or drafting questions for an appointment. Match the job to the tool.

Can AI replace my doctor?

No. It cannot examine you, order a test, prescribe, or take responsibility for the outcome. It also does not take a history unless you make it, which was one of the failure modes physicians identified. (1) Use it to arrive at your appointment better prepared, not to skip one.

Is ChatGPT reliable for medical advice?

For general education, often. For advice specific to you, treat every answer as a first draft. It generates the most statistically likely response, which is not the same as the correct response for your body, your history, and your medications. Verify anything you intend to act on.

Why would AI be less accurate for women in midlife?

Because it learned from research that under-represents you. Around 5% of clinical trials report results separately by sex, (3) and conditions including menopause, endometriosis and PMOS receive under 2% of private healthcare funding. (4) A model reflects the literature it read, so the thinner the literature on a topic, the more it defaults to the better-documented explanation.

What is the safest way to use an AI health coach?

Give it your own data rather than abstract symptoms, ask it what it would need to know about you to answer well, ask it to show its reasoning, and bring the output to a licensed clinician. Those four habits move it from an oracle you have to trust to a tool you can check.

Should I tell my doctor I used AI?

Yes. It tells them what you have been worried about and what you have already read, which makes twelve minutes go further. Anyone who makes you feel foolish for arriving informed is telling you something useful about the fit.

Maybe you just need better data

You do not need to become a data analyst to get more out of your own health information. You need to know which numbers matter and what each one is telling you.

That is what I put in the Lab Guide. It is free, it takes about fifteen minutes to read, and it is the thing I would hand you before you ask an AI health coach another question about your body.

Download the unBridled Longevity Blueprint™ Lab Guide →

This article is for informational and educational purposes only and is not medical advice. It is not intended to diagnose, treat, cure, or prevent any disease. Always consult your own healthcare provider.

References

  1. Draelos RL, Afreen S, Blasko B, et al. “Large language models provide unsafe answers to patient-posed medical questions.” npj Digital Medicine, published February 13, 2026. https://www.nature.com/articles/s41746-026-02428-5
  2. West Health–Gallup Center on Healthcare in America poll, conducted late 2025: approximately one quarter of US adults used an AI tool for health information or advice in the previous 30 days. Reported by PBS News, “Why so many Americans are using AI for health guidance.” https://www.pbs.org/newshour/health/why-so-americans-are-using-ai-for-health-guidance
  3. World Economic Forum, “The state of women’s health in numbers,” May 2026. https://www.weforum.org/stories/2026/05/womens-health-in-numbers/
  4. World Economic Forum, Women’s Health Investment Outlook 2026. https://www.weforum.org/publications/women-s-health-investment-outlook-2026/

CONNECT

elsewhere:

listen to the

podcast

subscribe to

newsletter

Subscribe to IN THE SADDLE: Smart news and wellness insights for midlife women who want to take the reins.


connect on

INSTA