AI in the Operating Room

From Surgical Performance Assessment to Real-Time Feedback

Abhirami Babu

Abhirami Babu

Director of Medical Affairs, SS Innovations International, Inc

More about Author

Abhirami Babu is a Junior Fellow of the International College of Surgeons and an Executive Council member of The Robotic Global Surgical Society (TROGSS), where she co-directs Robotic Timezone and leads Surgical Education and Global Partnerships at Surgical Excellence Academy. She serves as Director of Advocacy at Project IMG.

Every day, surgeons are judged by a metric no one can define. Surgical skill remains subjectively assessed, variably scored, and rarely measured in real time. AI is changing that, but only if calibrated against expert judgment. This article examines how validated AI makes surgical performance objective, improvable, and finally trustworthy.

The metric no one can define

Ask ten experienced surgeons to grade the same operative video, and you will often get ten different answers. For a profession built on precision, that is an uncomfortable truth: the way we assess surgical skill, the single biggest determinant of how a trainee progresses, how a department credentials, and how a patient fares, is still largely a matter of expert opinion.

Through The Robotic Global Surgical Society (TROGSS) and the Surgical Excellence Academy, where I work in advancing robotic surgery education and research, I see this from both ends: the talented trainee who lacks access to expert feedback and the program director who has no faculty hours to assess every case. Structured rating scales such as the Objective Structured Assessment of Technical Skill (OSATS) and, for robotic procedures, the Global Evaluative Assessment of Robotic Skills (GEARS) were real advances, but they still depend on a human expert watching and rating, which makes them slow, inconsistent, and impossible to scale. The cost is not abstract: it lengthens the learning curve and leaves quality leaders without the data to know who is truly ready to operate independently.

Why is this a global problem?

Robotic surgery is spreading far faster than the supply of expert assessors, and a society like TROGSS exists to close that gap, to make high-quality assessment portable across borders rather than confined to a few elite centers. A validated metric can reach a resident anywhere there is a camera and a connection, which makes objective assessment not just an efficiency gain for well-resourced hospitals but also a genuine equaliser for surgical training worldwide.

There is a structural problem we must be honest about as well. Today, much of robotic surgical training and credentialing is industry-led, designed, delivered, and certified by the companies that manufacture the platforms. That was understandable in the early years, but it ties a surgeon's competency record to a specific device and a commercial interest rather than to a portable, objective standard of skill. As the field matures, this has to change. We need platform-agnostic training and credentialing built around what a surgeon can actually do, measured the same way, whether they trained on one system or another, not around which console a vendor happened to sell their hospital. Objective AI assessment is what finally makes that possible, because the metrics it captures are properties of the surgery, not of the brand. We should be furthering surgical education with the same ambition we pour into furthering surgical technology. The hardware has raced ahead; the way we train and certify the humans using it has not kept pace.

Accreditation and privileging: the missing standard

This is no longer a niche concern for a handful of high-volume robotic centers. Minimally invasive surgery has spread well beyond its early home in general surgery, urology, and gynecology. Colorectal, thoracic, cardiac, hepatobiliary, bariatric, and head and neck surgeons now adopt laparoscopic and robotic approaches as a routine part of practice. As that shift reaches almost every surgical speciality, one question becomes central to surgical governance: how does a hospital decide that a given surgeon is competent to perform a given minimally invasive procedure?

Today, the answer is improvised and varies from one institution to the next. Privileges are often granted on the basis of a manufacturer's training course, a set number of proctored cases, and a colleague's sign-off. Those thresholds rest on case counts rather than demonstrated skill, and they lean on certifications designed by the companies selling the equipment. Advanced minimally invasive fellowships help, but there are too few of them, their quality varies, and most surgeons who take up these techniques never complete one. The result is a patchwork, with no shared definition of what competence in a procedure actually means and no portable record a surgeon can carry from one hospital to the next.

Objective assessment is what could finally anchor a common standard. If competence is measured the same way for every surgeon on any platform, a credentialing committee gains a defensible benchmark rather than a case count and a letter of recommendation. The standard would be set and owned by professional societies rather than device manufacturers, and a surgeon's competency record would travel with them. That is the pivot surgical education needs to make: from certification tied to a product to accreditation tied to skill.

What a camera can see that a clipboard cannot

Computer vision can take something we have always judged by feel, such as surgical dexterity, and make it measurable. Operating rooms are already saturated with signal: laparoscopic and robotic systems generate high-resolution video of every movement, and robotic platforms stream instrument kinematics, the position and trajectory of every tool tip thousands of times per second. Almost none of that record has been used for assessment.

From this data, systems can extract metrics that map onto what faculty have always looked for: economy of motion, which separates the deliberate hand of an expert from the redundant motion of a novice; tempo, including idle time and the rhythm of progression; tissue handling; and workflow recognition that flags when a step is performed out of order, skipped, or repeated. None of this is speculative. Public datasets such as JIGSAWS showed years ago that these features could classify surgeons by skill level with high accuracy. The real question now is whether we can measure well enough to trust a score for a decision that affects a trainee's career or a patient's safety.

EMMA and CARS: from raw metrics to meaningful competency

A path-length number is not, by itself, a competency judgment. The hard part is translating raw metrics into the structured competency domains that educators and credentialing committees actually use.

That is the problem we set out to solve. Working with my mentor, Professor Rodolfo J. Oviedo, and our computer-vision partners at VisionMed, we built two linked tools. CARS, our Comprehensive Assessment of Robotic Surgery framework, defines the domains an assessment should measure: tissue handling, efficiency, instrument control, bimanual dexterity, and operative judgment. EMMA, an Expert Multimodal Medical Agent, scores performance against those domains, ingesting video, kinematics, and procedural context to produce a domain-by-domain assessment in minutes.

The design philosophy matters as much as the architecture. We did not build EMMA to replace the surgeon-educator but to extend their reach. Because it scores the same constructs faculty already teach, its output is legible to a human expert and gradable against their judgment. A tool built by engineers alone optimises for what is easy to measure; one built with surgeons stays anchored to what matters in the field. For credentialing and privileging, an auditable, quantitative record is far more defensible than a recalled opinion.

The part most vendors will not mention: calibration

Here is where enthusiasm must meet rigor. A model that produces a confident-looking score is not the same as one that produces a correct score, and an assessor is only as trustworthy as the expert judgment it was calibrated against. Before any automated score influences a real decision, it must be validated for agreement with expert raters, tested across the full range of skill rather than only the easy extremes, and re-checked as it meets new surgeons, procedures, and platforms. A hospital that adopts a tool without asking "validated against whom, and how well do they agree?" is buying a number, not a measurement. Building that discipline from the start is what we are most committed to at TROGSS and the Academy.

From assessment to real-time feedback

Reliable measurement is only the foundation. The real goal is feedback, ideally delivered while it can still change the outcome. The most mature application today is the structured post-operative debrief an automatically generated, metric-anchored summary of how a case went, ready for the trainee and faculty to review together. This alone is transformative, because it turns every operation into a documented learning event instead of a memory that fades by the next morning.

Near-real-time feedback, flagging a deviation from expected workflow, an unusual instrument trajectory, or rising motion inefficiency during a case, is the focus of our ongoing VISION work on intraoperative assessment. It is technically within reach, but it carries a high bar: it must be accurate enough not to distract or mislead a surgeon at a critical moment, and it must support the operator's attention rather than fragment it

What hospitals should actually do?

For executives and quality leaders, the move is to treat these tools as clinical instruments, not gadgets. That means capturing operative video and kinematics securely, demanding validation evidence before adoption, governing how scores are generated and used, and tracking the regulatory and liability context around software as a medical device. Most importantly, use these tools to develop surgeons, not to surveil them. The fastest way to waste objective metrics is to weaponise them against trainees. Used well, they shorten learning curves and give every surgeon, not only those with a generous mentor, a fair and consistent path to mastery.

The road ahead

Objective assessment is the near-term prize, and it is achievable now. The longer arc points toward an autonomy assurance layer: the same validated measurement that tells us whether a human is operating well is what we will need to tell us whether an increasingly capable surgical system is doing so safely. The discipline we build now, calibration, concordance, and governance, is the groundwork for everything that follows.

For a century, we have asked surgeons to master a craft we could not measure. Applied with rigor rather than hype, this technology finally lets us measure it, and so teach it better, and teach it everywhere. That is the real promise: not to replace the surgeon's hands, but to make every pair of them, wherever they train, measurably better.

--Issue 73--