There is a version of AI in fertility care that sells well in a conference keynote and a version that survives contact with a clinic. They are not the same software, and the difference is mostly about which decisions the model is allowed to own.

Start with the decisions nobody wants to make

The highest-value applications in a fertility clinic are the ones clinicians are relieved to hand over, because they are not really clinical judgement at all — they are ranking and attention problems.

  • Who to call today. A counsellor with a caseload of a hundred and forty and time for fifteen conversations is making a triage decision whether or not they admit it. Ranking that list on disengagement risk is a clear improvement over ranking it on memory.
  • Which results need a human now. Flagging out-of-range values against the patient’s own history and protocol stage, rather than against a generic reference range.
  • Where the schedule is about to break. Predicting which branch-day will be short-staffed given roster, expected retrievals and historical no-show rates.
  • What is missing. Surfacing incomplete investigations, unsigned consents and un-actioned orders before they become a delay.

None of these produce a clinical recommendation. All of them change what a clinician sees first, which is where most of the practical benefit sits.

The clinical line, and why it should be bright

Embryo selection, protocol choice and dose adjustment are the applications everyone asks about. They are also where a model’s failure modes are least visible to the person acting on the output.

An embryo grading model that is ninety-four percent concordant with senior embryologists is genuinely impressive and genuinely unsafe to deploy silently, because the six percent is not randomly distributed — it concentrates in exactly the ambiguous cases where the model’s confidence is least warranted and the clinician’s attention is most needed.

The workable posture is narrow: the model may rank, annotate and explain. The named clinician selects, and the record stores both what was suggested and what was chosen. When those diverge, that is signal worth keeping, not an exception to suppress.

Four governance requirements worth insisting on

1. Every suggestion is attributable and stored

If a model influenced a decision, the record should say so — which model, which version, what inputs, what it said. A recommendation that vanishes after it is displayed cannot be audited, and an unauditable recommendation should not be shown in a clinical setting.

2. The clinician can always see why

Not a saliency map for its own sake, but the concrete inputs: this couple is ranked high because there has been no contact for nineteen days, two investigations are outstanding and an appointment was cancelled. That is a reason a counsellor can act on or dismiss on the spot.

3. Overrides are cheap and recorded

If disagreeing with the model is slower than complying with it, you have built an automation bias machine. Overriding should take one click and should be a first-class event in the record, not a workaround.

4. Performance is monitored against your population

A model trained elsewhere has a distribution shift problem the day it arrives. Concordance should be measured continuously against your own outcomes, by branch, and the results should be visible to the clinical lead rather than to the vendor alone.

The question to ask a vendor

Not “what is the accuracy?” — every answer to that is true and uninformative. Ask instead: when the model is wrong, how will we find out? If the answer involves the clinic noticing a pattern in its own outcomes months later, the governance is not finished.

Built into FertilityNXT

Traceable lab handling and versioned consent are part of the platform. Start a free trial to see them against your own workflow.