Putting Google’s Hypothesis Generation to the test
A patient recently sat in my office and shared a story about steak and tick bites. She had a tick bite months earlier, but there was nothing remarkable at the time. Now, hours after eating red meat, she developed hives and bowel issues. The diagnosis was alpha-gal syndrome, a tick bite-acquired IgE allergy to a sugar molecule, galactose-α-1,3-galactose, found in mammalian tissue.
Every day, I see a handful of patients affected by alpha-gal. Some cannot even take their regular medications because they contain substances derived from meat. The treatment I offered was simple: Don’t eat meat because there is no drug that modifies the disease. However, my patients don’t ask, “What should I avoid?" Instead, they ask, “Can this be fixed?”
I put that question to Google’s Hypothesis Generation tool in Gemini for Science. It is an experimental artificial intelligence (AI) co-scientist that generates hypotheses, reviews its own work, ranks them through pairwise debate, and iterates in a self-improving loop.
I'll tell you what it gave me—and what it couldn't.
The most consequential step was not the prompt. I didn't write a big prompt. What actually happened was an interview. I stated the problem, and the system started asking questions: What counts as therapeutic success? Which drug classes are acceptable? What route? What timeframe? What cost? That exchange forced me to commit to an endpoint that a patient would care about— not a lower IgE titer, which means nothing at the dinner table, but durable, drug-free tolerance, good enough to pass an oral food challenge with a normal meal of mammalian protein.
Then I set the constraints: oral agents preferred, three to six months of treatment, under $100 a month, adults aged 18 to 65, and no broad immunosuppressants or biologics. Given those specifications, the system mapped three solution landscapes: selective IgE modulation, gut barrier and antigen processing, and tolerance reprogramming. It came back with 10 ranked directions, each with a mechanistic rationale, a protocol sketch, and a formal evidence assessment.
The top proposal was a metformin-guided, phased oral immunotherapy. The reasoning: metformin activates AMPK, suppresses the lipid-driven unfolded protein response in enterocytes, and stabilizes tight junctions, which opens a non-inflammatory window. During that window you escalate the antigen from low-fat glycoproteins to mammalian glycolipids. Other directions layered zileuton, ezetimibe, ursodiol, oral cromolyn, montelukast, pioglitazone, or calcitriol onto the same scaffold.
The system also flagged potential issues with the generated hypotheses. A “chylomicron-suppression paradox” suggested that metformin’s own suppression of apoB-48 secretion could blunt the protective IgG4 response on which the protocol depends. In addition, metformin’s GI side effects could mimic the GI-predominant alpha-gal syndrome phenotype, so you’d struggle to tell the difference between drug effect and symptoms of the disease. Many drugs also contain trace amounts of animal protein that could trigger an allergic response.
Now for the caveats. These are hypotheses, not evidence. No laboratory, animal, or clinical validation was performed, and nothing here is clinical guidance. The strongest human signal—alpha-gal syndrome remission in bariatric patients who happened to be on metformin by itself—is not solid evidence. The barrier-to-tolerance mechanism rests largely on murine and in vitro data. And some of the specific links are frankly speculative.
There’s also a detail that’s easy to miss. Excipient safety is a design constraint. Standard generics may carry mammalian-derived fillers with trace alpha-gal. And because these drugs are cheap and off-patent, premature self-experimentation by a motivated patient community is a real hazard. Any actual test requires institutional review board oversight and anaphylaxis-ready monitoring.
Hypothesis Generation also has its own limits. Generative systems can misattribute and fabricate references. The internal Elo rankings are automated evaluations with no external validity. I benchmarked nothing. There is no ground truth, no expert-generated comparison set, and no alternative model, so I can't honestly quantify novelty or accuracy. (A fuller account of the methodology, results, and limitations is available in a preprint.)
If you have an under-researched condition in your specialty, such as refractory hypoglycemia, rare monogenic diabetes, and nonsurgical treatment for primary hyperparathyroidism, you could try Hypothesis Generation. My advice is to interview the system rather than prompt it with a fully formed problem and to make it commit to a hard clinical endpoint before any hypothesis is generated. Encode your real constraints: route, duration, cost, population, comorbidities, and formulary.
This workflow can also convert a literature summary into a candidate protocol, so ask explicitly for failure modes, fragility assessments, and disconfirming experiments—and read those first. Verify every citation and appraise the output as you would any medical literature: plausibility, safety, and whether the endpoint is measurable. Google’s Hypothesis Generation is a useful tool for producing multiple testable hypotheses and rapidly surveying the literature behind each one. But use it as a starting point. While it can compress weeks of synthesis into an afternoon, the tool cannot compress the validation that must follow.
AACE Endocrine AI is published by Conexiant under a license arrangement with the American Association of Clinical Endocrinology, Inc. (AACE®). The ideas and opinions expressed in AACE Endocrine AI do not necessarily reflect those of Conexiant or AACE. For more information, see Policies.