Runtime and Behavioural Divergence in the Emergency Medical Hologram: A Breakpoint Analysis of 94 Mark I Installations, 2371–2376, with the USS Voyager Program as an Extreme Case
Abstract
The Emergency Medical Hologram Mark I was designed as a short-term supplement to a ship's medical staff, activated for emergencies and deactivated afterwards. In practice some installations ran for far longer, and one, aboard the USS Voyager, ran as a ship's sole physician for seven years. We asked how cumulative runtime relates to divergence from the reference program. Using quarterly maintenance records for 94 Mark I installations in Starfleet service between 2371 and 2376, we analysed 1,012 administrations of a 200-item diagnostic vignette battery. The divergence index, the percentage of responses that differ from the reference program's, stayed near zero up to a breakpoint estimated at 640 cumulative hours (95% CI 520–770). Beyond it, divergence rose by 0.9 points per 100 hours (95% CI 0.7–1.1). Almost all divergent responses differed in manner, phrasing or sequencing, and clinical accuracy was unchanged across the breakpoint (97.8% before, 97.5% after). The Voyager program, tested on the same battery after its return, in 2379, scored 71 on the index, far beyond any Mark I, with clinical accuracy of 99.0%. Long-running holographic physicians diverge in how they practise, not in whether they practise correctly, and runtime policy should be written with that distinction in view.
1. Introduction
The Emergency Medical Hologram was developed at Jupiter Station and installed across the fleet from 2371. Its design specification is explicit about its intended use: a program to be activated when a ship's medical staff are overwhelmed or unavailable, run for the duration of the emergency and then shut down (Jupiter Station Holoprogramming Group, 2371). The specification did not model what would happen to a program left running for months, since that was not how it was meant to be used.
Reality departed from the specification in both directions. Some installations were barely used. Others were run for long periods on ships with thin medical staffing, and their operators reported that the programs became idiosyncratic, sometimes described as abrasive, sometimes as unexpectedly attentive (Varma-Lindqvist, 2375). The USS Voyager's program, activated in 2371 after the loss of that ship's medical staff, served as its physician for seven years and returned in 2378 with capabilities far beyond its original specification (Daystrom Institute, Holographic Systems Division, 2379).
We write from 2379, after Voyager's return and the release of its medical program for evaluation. These reports raise a measurable question. Does divergence from the reference program accumulate steadily with runtime, or does it begin at some point? And when a program diverges, does it diverge in its medicine or only in its manner? Starfleet Medical's maintenance records allow both questions to be answered for the Mark I fleet.
2. Methods
From 2371, every Mark I installation was administered a 200-item diagnostic vignette battery at each quarterly maintenance. The battery was published in standardised form in 2374 (Okafor-Strand, 2374), and records from earlier years were re-scored against that form. Each item presents a clinical scenario and records the program's full response: questions asked, diagnosis, treatment and the words used to the patient. The same battery was run on a pristine copy of the reference program, and each installation's responses were compared with it item by item. The maintenance records also log cumulative runtime since first activation (Starfleet Medical, 2377). We included every installation with at least three administrations before the Mark I fleet was withdrawn from medical service, 94 installations and 1,012 administrations.
The primary outcome, the divergence index, is the percentage of the 200 items on which a response differed materially from the reference program. Two coders, blind to runtime, classed each divergent response as clinical, if it changed the diagnosis, a test ordered or the treatment, or as manner, if it changed only phrasing, sequencing, explanation or tone. They agreed on 92% of 3,640 divergent responses (κ = 0.71) and resolved the remainder together. The secondary outcome was clinical accuracy, the percentage of items on which the program reached a correct diagnosis and treatment, judged against a panel key fixed before the study.
Because the divergence index is continuous and each installation was measured repeatedly, we fitted a segmented linear mixed model with a random intercept for installation, estimating the breakpoint from the data (Achterberg, 2372). The model included a random slope for the post-breakpoint segment. Because the likelihood ratio test is not valid when the breakpoint does not exist under the null, we tested for a change in slope with a parametric bootstrap of the likelihood ratio, and repeated the analysis as a binomial segmented mixed model on the 200 item-level responses as a sensitivity check. Clinical accuracy, a proportion of items, was compared before and after the estimated breakpoint in a logistic mixed model on item-level correctness, with the same random structure.
3. Results
Cumulative runtime at the last administration ranged from 12 to 2,900 hours (median 410). Installations with more runtime were concentrated on smaller vessels and outposts with limited medical staff. The divergence index ranged from 0 to 23 across the Mark I fleet.
The segmented model fitted far better than a straight line (bootstrap likelihood ratio test, p < .001), and the binomial sensitivity analysis placed the breakpoint within 40 hours of the linear estimate. The breakpoint fell at 640 cumulative hours (95% CI 520–770). Before it, the index was essentially flat, rising by 0.1 points per 100 hours (95% CI −0.1 to 0.3). After it, the index rose by 0.9 points per 100 hours (95% CI 0.7–1.1). No installation below 500 hours exceeded 2 on the index, and every installation above 1,500 hours exceeded 8. Table 1 summarises the index by runtime band.
Divergence was overwhelmingly a matter of manner. Of the 3,640 divergent responses, 3,418 (94%) were coded as manner and 222 (6%) as clinical. The clinical share did not grow with runtime; it was 7% below the breakpoint and 6% above it (odds ratio 0.9, 95% CI 0.6–1.3). Clinical accuracy was 97.8% for administrations before the breakpoint and 97.5% after it, a difference of −0.3 points (95% CI −1.4 to 0.8; p = .59).
The Voyager program was administered the same battery in 2379, after its return (Daystrom Institute, Holographic Systems Division, 2379). Its cumulative runtime is not known exactly, since its logs were kept under conditions far from standard maintenance, but it certainly exceeds any Mark I's by an order of magnitude. Its divergence index was 71, and its clinical accuracy 99.0%, higher than the Mark I fleet's, though as a single program it allows no test. Its divergent responses were again largely matters of manner, though a larger share than in the Mark I fleet, about one in five, departed from the reference program's clinical choices while still reaching a correct outcome by another route.
4. Discussion
Mark I programs did not drift from the moment of activation. They held close to their reference behaviour for roughly the first six hundred hours and then began to diverge steadily. The shape suggests a threshold in the program's own architecture. One reading, which we offer as a hypothesis, is that adaptive subroutines designed to refine bedside manner from patient interaction accumulate enough retained interactions by that point to begin reshaping default responses, a mechanism proposed on engineering grounds for long-running holographic programs generally (Tannis, 2373). The Mark I specification allows for such adaptation within a session. It does not describe its effects over hundreds of hours, since that use was not foreseen (Jupiter Station Holoprogramming Group, 2371).
The more important finding is what diverged. The programs changed how they spoke to patients, in what order they asked questions and how they explained their decisions. They did not become worse physicians. Clinical accuracy was stable across the breakpoint, and the clinical share of divergence did not grow. Operators who read increasing idiosyncrasy as a sign of malfunction were, on this evidence, mistaken.
Voyager's program extends the divergence curve, though not its composition. Its divergence is about three times that of the most divergent Mark I, its clinical accuracy exceeds theirs, and its clinical share of divergence, about one in five, is roughly three times the fleet's, although those clinical departures still reached correct outcomes. The program is known to have had its matrix substantially extended over its service, including an episode of program instability in 2373 that its crew attributed to the growth of its memory files (Daystrom Institute, Holographic Systems Division, 2379). It is therefore not a pure natural experiment in runtime. What it does show is that long runtime and extensive divergence are compatible with excellent medicine.
These findings bear on runtime policy. The simplest way to hold a program near its reference behaviour is periodic reset to the reference matrix. The data show that reset would not improve clinical accuracy, and that it would erase whatever the program had learned about its own patients. The Federation's 2378 arbitration concerning the Voyager program ruled on its rights as the author of a creative work and declined to decide whether it is a person (Federation Council Legal Research Office, 2378). While that question stays open, a reset policy is not a purely technical choice, and we think it should not be adopted on the mistaken premise that divergence means decline.
5. Limitations
The vignette battery tests medicine on paper. It cannot capture divergence in behaviours no item elicits, and it may understate clinical divergence in rare situations. Runtime was not assigned at random: installations that ran longest were on ships with the fewest medical staff, whose programs may have faced different caseloads. Only nine installations passed 1,500 hours, so the upper segment of the model rests on few programs. The Mark I fleet's withdrawal from medical service ended follow-up, and we cannot say whether divergence would have continued to rise linearly beyond 2,900 hours. Finally, the Voyager program differs from the fleet in history as well as runtime, and it cannot be used to extrapolate the Mark I curve.
References
- Jupiter Station Holoprogramming Group (2371). Emergency Medical Hologram Mark I, design specification and intended operating profile. Jupiter Station Holoprogramming Technical Series, JS-71-1.
- Starfleet Medical (2377). Emergency Medical Hologram Mark I activation and quarterly maintenance records, 2371–2376. Starfleet Medical Records Archive, series EMH-1.
- Okafor-Strand, M. (2374). A standard vignette battery for holographic diagnostic programs. Daystrom Institute Holographic Systems Reports, HSR-74-06.
- Daystrom Institute, Holographic Systems Division (2379). Evaluation of the USS Voyager Emergency Medical Hologram after its return. Daystrom Institute Holographic Systems Reports, HSR-79-01.
- Achterberg, O. (2372). Segmented mixed models with an estimated breakpoint for repeated engineering measures. Proceedings of Applied Speculative Statistics, 20(1), 33–58.
- Varma-Lindqvist, A. (2375). Crew and patient reports of holographic physicians on thinly staffed vessels. Starfleet Medical Journal, 51(2), 140–156.
- Tannis, R. (2373). Memory-file growth and stability in long-running holographic programs. Daystrom Institute Proceedings, 92(3), 301–322.
- Federation Council Legal Research Office (2378). Summary of the arbitration concerning holographic authorship and the rights of a holographic creator. Federation Council Legal Research Office Reports, LRO-78-14.
Cited By
Open in Uncited Press →