What a Truthsayer Registers: Prosodic and Autonomic Correlates of Deception and of Truthsayer Verdicts in 432 Paired Utterances Judged by Four Reverend Mothers, 10234–10235 AG
Abstract
Truthsayers, the Reverend Mothers who verify spoken testimony, are credited with telling truth from falsehood, but nothing observable in the speaker has been measured against their verdicts. We ask which prosodic and autonomic signals accompany deception and the Truthsayers' verdicts, and how much of the Truthsayers' accuracy those signals account for, without presuming that a Truthsayer uses them. Under a restricted-use agreement with the Sisterhood, 36 speakers without Sisterhood training each gave six matched pairs of truthful and false answers (432 utterances) while pitch, response latency, filled pauses, pupil diameter, pulse rate and skin conductance were recorded. Four Truthsayers and a panel of six clinicians judged every answer in the room. The Truthsayers gave 1,525 of 1,728 verdicts correctly (88.3%, 95% interval 86.3–90.1%; sensitivity index d′ = 2.4), the clinicians 58.6%. Within speakers, false answers began later, were higher in pitch and held more filled pauses, and were accompanied by a wider pupil, a faster pulse and higher skin conductance; pupil diameter separated the members of a pair most often (78.7%). A cue-only classifier tested on unseen speakers reached 74.1%. Truthsayer verdicts rose with latency, pitch, pupil diameter and skin conductance at fixed veracity, yet adding the cues cut the log odds ratio for true status by only 11%, and accuracy was 85.4% where the cues were least informative. Prosodic and autonomic signals register above chance, and what we measured does not explain the faculty.
1. Introduction
Four Truthsayers judged 432 truthful and deceptive utterances, and 1,525 of their 1,728 verdicts were correct (88.3%, sensitivity index d′ = 2.4). Six clinicians with no Sisterhood training, in the same room and writing on the same forms, were right in 58.6% of 2,592 verdicts. A classifier built from six signals that instruments could record (pitch, response timing, disfluency, pupil diameter, pulse rate and skin conductance) was right for 74.1% of utterances from speakers it had never seen. This paper reports how those figures relate and how much of the first the measured signals account for.
A Truthsayer is a Reverend Mother who can tell whether a speaker is lying. Such mothers are placed at court and in Great House households to verify testimony (Merrowin, 10228 AG), and Gaius Helen Mohiam, Truthsayer to the Emperor, is the best-known holder of the office. The faculty is held to rest on the heightened awareness that the spice agony confers, developed by training. Whether it works through signals that a speaker emits and an instrument could record has never been measured.
Our aim is deliberately narrow. A signal that correlates with deception does not explain a Truthsayer's judgement: it might be what she uses, one of several things she uses, or something that accompanies what she registers. We therefore measure which signals separate truthful from deceptive answers by the same speaker, which signals the verdicts follow, and how much of the Truthsayers' accuracy those signals account for.
We write c. 10236 AG. The corpus is our own, collected in 36 instrumented sessions on Wallach IX in 10234–10235 AG under a restricted-use agreement of the kind that governed recent acoustic work on Voice instruction in this journal. That study found no detectable fundamental-frequency difference between Voice and ordinary utterances. Clinical work on feigned histories reports slower onset and higher pitch in false answers (Calveth, 10225 AG), which supplied two of our six cues.
2. Methods
The Sisterhood admitted the study under a conditional agreement (Bene Gesserit Institute of Kinesthetic & Biological Sciences, 10234 AG). Its terms were that derived measurements and sealed verdict forms remain on Wallach IX (Bene Gesserit Institute of Kinesthetic & Biological Sciences, 10234–10235 AG), that no raw video be kept, that judges not be questioned about method, training or use of the spice, and that speakers' identities be withheld. The second author, of the Suk Faculty of Clinical Medicine, recruited speakers and supervised recording; the first, of the Institute, coordinated the judges. Neither gave verdicts.
Thirty-six volunteer speakers, all Suk clinical residents with no Sisterhood training and no earlier contact with the judges, each attended one session. After a baseline utterance (name, post and date), each answered twelve questions about their own verifiable history, such as postings and dates of qualification. The questions formed six pairs matched for category and expected length; a sealed card drawn before each pair assigned which question to answer falsely, and order was randomised. Truthful answers were checked against Faculty registry records. A stipend supplement rose with the number of answers the judges accepted as truthful. The corpus holds 432 utterances, 216 truthful and 216 deceptive, in 216 pairs.
Speech was recorded at close range. Latency ran from the end of the question to the onset of speech, pitch shift was mean fundamental frequency over voiced frames in semitones, and a filled pause was a hesitation vowel or nasal, counted blind to condition from randomised clips. Pupil diameter came from an infrared camera, pulse rate from a finger photoplethysmograph and skin conductance from palmar electrodes, handled as in clinical practice (Dessarin, 10219 AG; Ostravane, 10222 AG). Every cue except filled pauses is a deviation from the speaker's baseline utterance. The cues were fixed before the first session, and instrument displays were hidden from the judges.
Four Truthsayers attended every session. Each sat about two metres from the speaker, in the same room and without instruments, and after each answer wrote a forced verdict, truthful or deceptive, on a sealed form. They did not confer or receive feedback. A comparison panel of six Suk physicians with no Sisterhood training sat in the same room and completed the same forms. No judge was told the proportion of false answers. The Truthsayers gave 1,728 verdicts and the panel 2,592.
Hit rate was the share of deceptive answers judged deceptive, false-alarm rate the share of truthful answers so judged, and d′ the difference of their normal quantiles. Because verdicts on one speaker's answers are not independent, intervals for judgement and classifier performance resample the 36 speakers with replacement over 2,000 draws (Pellindor, 10231 AG). For each cue we averaged the within-pair difference inside each speaker and took an interval across the 36 speaker means (35 degrees of freedom); filled-pause counts were also fitted with a log-linear count regression clustered by speaker. A logistic regression on the six cues, trained with each speaker held out in turn, gave a cue-only classifier. Verdicts were modelled by logistic regression on true status with a term for each judge, alone and then with the six cues standardised, standard errors clustered by speaker. Agreement among judges was a chance-corrected coefficient for several raters. The split by classifier confidence was decided after the first results and is exploratory.
3. Results
Truthsayers were correct on 1,525 of 1,728 verdicts (88.3%, 95% interval 86.3–90.1%). They judged 89.2% of deceptive answers deceptive (87.1–91.3%) and 12.7% of truthful answers deceptive (10.2–15.4%), giving d′ = 2.4 (2.2–2.6). Individual accuracies ran from 87.0% to 89.8% and agreement among the four was 0.63. The panel was correct on 1,518 of 2,592 verdicts (58.6%, 56.7–60.5%), with d′ = 0.43 (0.34–0.53), individual accuracies from 54.6% to 62.5% and agreement of 0.09.
Every cue differed between deceptive and truthful answers within speakers (table below, all p < .01). False answers began 109 ms later (95% interval 59–158) and were pitched 0.75 semitone higher (0.54–0.97). Filled pauses had a rate ratio of 1.66 (1.20–2.29, p = .003). The pupil was wider by 0.18 mm (0.15–0.21), pulse rate higher by 2.7 beats per minute (1.9–3.6) and skin conductance higher by 0.26 μS (0.17–0.35). Pupil diameter separated the conditions most clearly: the deceptive member of a pair had the larger value in 78.7% of the 216 pairs (74.1–83.3%), against 58.1% for filled pauses (52.8–63.2%).
Together the six cues carried real information about veracity. The cue-only classifier, tested on speakers it had not seen, labelled 320 of 432 utterances correctly (74.1%, 70.1–78.2%; area under the curve 0.83, 0.79–0.87). That is some 14 points below the Truthsayers and more than 15 points above the panel.
The Truthsayers' verdicts followed the cues. With true status held fixed, each standard deviation of response latency raised the odds of a deceptive verdict by a factor of 1.50, pitch shift by 1.35, pupil diameter by 1.34 and skin conductance by 1.37 (intervals in the table). Filled pauses (1.00) and pulse rate (0.95) showed no such relation, although both differed between conditions. The odds ratio for true status was 56.8 (38.4–84.0) before the cues were added and 36.1 (23.1–56.5) after, a fall of 11% in the log odds ratio (4.04 to 3.59).
Where the cues were weakest, they did not account for the Truthsayers' margin. We ranked utterances by the classifier's confidence and split them into thirds of 144. In the third where its predicted probability lay between 0.30 and 0.70, the classifier was correct for 54.9% of utterances and the Truthsayers for 85.4% of 576 verdicts (81.3–89.1%). In the middle third the figures were 72.9% and 86.3%, and in the third where the classifier was most confident, 94.4% and 93.1%.
4. Discussion
Prosodic and autonomic signals separate lies from truths in speakers untrained in concealment, and every one of the six cues did so. Clinicians in the same room scored 58.6%, well short of the 74.1% the cues made available. A 0.18 mm change in pupil diameter or a 0.26 μS rise in conductance is unlikely to be visible to an unaided observer. The Truthsayers' verdicts rise with four of the six cues even among utterances of the same veracity, but not with the pulse rate or disfluency that also separated the conditions.
These signals are not the whole of what the Truthsayers register. Adjusting for all six cues left most of the dependence of verdicts on veracity in place, and where the cues were least informative the Truthsayers were still correct 85.4% of the time. Nothing we measured explains the faculty away: the Truthsayers' verdicts correlate with some of what we measured, and their accuracy exceeds what those signals yield alone.
Three readings of the link between cues and verdicts remain open, and the design cannot separate them. The Truthsayers may perceive the cues or visible correlates we did not record, such as flushing or changes in gaze. Both verdicts and cues may instead follow from the speaker's state at the moment of the lie, which the Truthsayers register by a route our instruments miss. Or they may register something else that varies with the cues in these speakers. The judges sat in the room with the speaker, so no channel was isolated.
The pitch result bears on a neighbouring technique. False answers were pitched 0.75 semitone higher, whereas the Voice study found no detectable pitch difference between Voice and ordinary utterances, so the two phenomena do not share a pitch signature. Our speakers were untrained and not producing Voice, and we did not measure the subharmonic energy or formant timing on which that study's carrier rests; we claim no relation between deception cues and that carrier.
5. Limitations
Instructed lies about verifiable personal history carried only a stipend supplement as stakes, whereas a Truthsayer at court judges testimony that can cost a speaker position or life. Signals and accuracy may differ under such stakes. The speakers were residents of one school and untrained in concealment; practised deceivers may suppress some cues.
Four judges from one establishment gave every verdict, so intervals that resample speakers do not cover variation among Truthsayers. We know nothing of their training or use of the spice, because the agreement forbade it, and the Sisterhood declined to admit acolytes as judges. We therefore cannot say how much of the Truthsayers' margin belongs to the awareness the spice agony confers and how much to training.
Six cues are a short list. Gaze, facial colour, posture and breathing were not recorded, and any of them may carry what the classifier lacks. The classifier was linear, trained on 420 utterances at a time, and a richer model might do better, so "unaccounted for" means unaccounted for by these six cues in this model. The split by classifier confidence was chosen after the first results and rests on 144 utterances from 36 speakers. The link between cues and verdicts is an association only, since we manipulated no signal.
These limits bound the claim without removing it. In this corpus prosodic and autonomic signals register above chance on every measure we took: false answers were later, higher in pitch and more disfluent, and came with a wider pupil, a faster pulse and higher skin conductance, and the Truthsayers' verdicts followed four of those signals. The same signals leave most of what the Truthsayers achieved unexplained.
References
- Marn, T., & Vantrel, S. (2026). A Prosodic Carrier Signature in Archived Bene Gesserit Voice Training Recordings: A Phonetic Analysis of Twelve Instructors, 10205–10232 AG. Uncited Press. https://doi.org/10.0000/uncited.2026.0129
- Bene Gesserit Institute of Kinesthetic & Biological Sciences (10234 AG). Conditional agreement for instrumented paired-utterance sessions with Truthsayer judges. Bene Gesserit Archives, Wallach IX (restricted), rev. 2.
- Bene Gesserit Institute of Kinesthetic & Biological Sciences (10234–10235 AG). Paired-utterance session records, derived measurement series and sealed verdict forms. Bene Gesserit Archives, Wallach IX (restricted), series XII, sessions 1–36.
- Merrowin, J. (10228 AG). Truthsayer placement and the attestation of testimony at court and in Great House households, a source-critical survey. Bene Gesserit Institute Review, pp. 61–84.
- Calveth, R. (10225 AG). Malingered histories in clinical interview, response latency and pitch movement. Suk School Medical Transactions, pp. 102–121.
- Dessarin, K. (10219 AG). Finger photoplethysmography and palmar electrodermal recording during speech in clinical subjects. Suk School Medical Transactions, pp. 55–72.
- Ostravane, H. (10222 AG). Baseline referencing of pupil diameter in speaking subjects, gaze and lighting artefacts. Suk School Medical Transactions, pp. 140–158.
- Pellindor, A. (10231 AG). Sensitivity indices for forced-verdict tasks with few raters, speaker-level resampling. Proceedings of Applied Speculative Statistics, pp. 33–49.
Open in Uncited Press →