Uncited Press Open the interactive journal →
Star Wars · Linguistics & Semiotics

Where Protocol Droids Mistranslate in Huttese-Basic Trade Negotiation: A Corpus Study of Error Types and Their Predictors, 29–25 BBY

Dr. Coraline Vess1, Dr. Arden Kolvari1
1 Coruscant Academy of Letters and Sciences, Faculty of Languages
Received 5 Aug 2026 · Revised 3 Sep 2026 · Accepted 16 Sep 2026 · DOI: 10.0000/uncited.2026.0824

Abstract

Protocol droids routinely interpret between Huttese and Galactic Basic in Outer Rim trade, yet the errors in their renderings have not been classified or explained. We asked which properties of an exchange predict a translation error. The Faculty of Languages at the Coruscant Academy of Letters and Sciences holds 91 recorded and transcribed negotiation sessions. After exclusions we annotated 1,967 exchanges from 72 sessions recorded between 29 and 25 BBY. Annotators coded each Huttese-to-Basic rendering for four error types: register or honorific mismatch, idiom rendered literally, quantity or price ambiguity, and omission. An error occurred in 343 exchanges (17.4%). In a logistic model with standard errors clustered by session, the odds of error were 2.78 times higher when an idiom was present (95% CI 2.19–3.53), 1.58 times higher in colloquial than in formal speech (1.25–2.00), 1.21 times higher per standard deviation of speech rate (1.07–1.37) and 1.36 times higher per ten turns elapsed since a glossary term was set (1.17–1.57). With idiom-type errors removed from the outcome, the idiom odds ratio was 1.48 (1.14–1.92). The last result is compatible with a weakening hold on session-specific terms as a negotiation proceeds, although position within the session cannot be separated from it in these data. We propose that idiom flagging and periodic glossary refresh be tested as interventions in interpreted negotiation.

1. Introduction

Huttese is the language of the Hutts and is widely spoken across the trade and underworld networks of the Outer Rim, where syndicates dominated by the Hutts conduct much of their business. Galactic Basic is the common tongue of the wider galaxy. A buyer from the Core Worlds who wishes to deal with a Huttese-speaking seller therefore commonly works through an interpreter, and that interpreter is very often a protocol droid, a class of machine described as fluent in vast numbers of languages.

Fluency is a statement about repertoire. It says little about fidelity in a live negotiation, where a misrendered price or a flattened courtesy can alter a bargain. Earlier work on interpreted trade talk has described what is lost in omission (Duvane, 27 BBY), in price and quantity terms (Orlun & Varsk, 29 BBY) and in the marking of register (Tessane, 28 BBY), and has treated speech rate as a source of rendering loss (Ossel, 25 BBY). Bench evaluations of context handling in translation droids have examined how well a droid holds on to terms that a session has fixed (Pellard, 26 BBY). Most of these studies rely on constructed test speech or on a handful of sessions. None has set several candidate predictors against one another in recorded negotiation.

This paper is written from the vantage of 24 BBY and draws on a corpus of recorded and transcribed negotiation sessions held by the Faculty of Languages of the Coruscant Academy of Letters and Sciences. We had two aims. The first was to propose a coding scheme for translation errors in droid-mediated Huttese-Basic negotiation. The second was to test which properties of an exchange predict an error, taking four from the literature: the presence of an idiom, colloquial as against formal Huttese, the speaker's rate of speech, and the number of turns since a glossary term was fixed. The last is an operational hypothesis about behaviour, namely that an interpreter's rendering of a session-specific term becomes less reliable as turns accumulate. We make no claim about how any droid stores or retrieves the content of a session.

2. Methods

Corpus. The recordings were made between 29 and 25 BBY in trade halls on several Outer Rim worlds, of the kind surveyed by Teshnar (30 BBY), and were transcribed by Faculty staff in Huttese and in the Basic rendering that the interpreter actually produced. The Faculty holds 91 sessions. We excluded 19: eight because the audio was too poor to transcribe reliably, six because the parties moved into Basic without the interpreter for a substantial part of the session, and five because the interpreting droid was replaced partway through. That left 72 sessions, each interpreted by a single protocol droid. Analysis was completed in 24 BBY.

Handling of recordings. The recordings were made by hall operators and released to the Faculty on condition that parties and goods be anonymised. Names, cargo descriptions and prices were masked in the working transcripts, and no participant can be identified from this paper. Interpreting droids were identified by session only. We had no access to their make, configuration or service history, so nothing here concerns differences between droids.

Exchange unit. An exchange is one Huttese speaking turn together with the interpreter's Basic rendering of it. From each session an annotator selected a contiguous stretch of 22 to 34 exchanges (mean 27.3), beginning within the first five exchanges after the opening business of the session, which produced 1,967 exchanges in all. Only the Huttese-to-Basic direction was annotated, because the Basic-to-Huttese direction could not be checked against the transcripts with the same confidence.

Error coding. An exchange was coded as containing an error when the rendering altered, obscured or dropped content that a fluent bilingual reader of the transcript judged material to the negotiation. Four types were defined for this study. A register or honorific mismatch is a rendering that conveys a different level of deference, or a different form of address, from the source, whether by flattening ceremonial address into neutral speech or by inflating a plain remark. An idiom rendered literally is a fixed expression translated word for word, so that the Basic version has a meaning the speaker did not intend. A quantity or price ambiguity is a rendering that leaves a number, a unit or a price term open to more than one reading, or shifts it. An omission is a rendering that leaves out a clause or term present in the source. When an exchange showed more than one error, it was coded to the single most consequential, so that each erroneous exchange carries exactly one type.

Predictors. Register was coded as formal or colloquial from address forms, ritual openings and lexical choice, using guidelines we developed for the purpose. An exchange was coded as containing an idiom when it included an expression from the working inventory of Huttese fixed expressions (Vess, 27 BBY) or one that two annotators agreed was non-compositional. Speech rate was measured in syllables per second from time-aligned transcripts; across exchanges it averaged about 4.6 (SD about 0.9). Turns since a glossary term was set was counted from the most recent exchange in which the parties fixed a term, such as a unit of account, a cargo grade or a form of address, and the interpreter rendered it. The count began at the session's opening exchange of terms, so the annotated stretches start with values from 0 to 4, and the largest value observed was 37.

Reliability. Each exchange was coded by one annotator. A random sample of 300 exchanges was coded independently by a second. Both flagged an error in 50 exchanges, neither did in 232, and they disagreed on 18, giving Cohen's κ = 0.81 (95% CI 0.73–0.90) for the presence of an error. Among the 50 exchanges flagged by both, the two annotators assigned the same type in 43 (86%). Disagreements were resolved with a third Faculty reader, and the resolved codes were used. On the same sample, agreement was κ = 0.87 (95% CI 0.81–0.94) for the coding of idiom, with both annotators flagging 66 exchanges, neither flagging 220 and the two disagreeing on 14, and κ = 0.83 (0.76–0.89) for register, with 128 exchanges coded colloquial by both, 146 formal by both and 26 disagreements.

Analysis. The outcome is binary at the level of the exchange, and exchanges within a session share an interpreter, a pair of parties and a hall, so we fitted a logistic regression with standard errors clustered by session, applying the small-sample factor for 72 clusters. Register and idiom were entered as binary terms, speech rate per standard deviation, and turns since a glossary term was set per ten turns, with the last two assumed linear on the log-odds scale. We report unadjusted and mutually adjusted odds ratios with Wald 95% confidence intervals and a normal reference distribution, and predicted probabilities at stated covariate values. The four predictors were chosen from the earlier literature and no correction was made for multiple testing. Because an idiom-type error can occur only in an exchange containing an idiom, so that the idiom predictor is partly tied to the outcome by definition, we repeated the model as a sensitivity analysis with idiom-type errors removed from the outcome.

3. Results

An error was coded in 343 of the 1,967 exchanges (17.4%). Omission was the commonest type (99 errors, 28.9% of all errors), followed closely by register or honorific mismatch (98, 28.6%) and quantity or price ambiguity (90, 26.2%). Idiom rendered literally accounted for 56 errors (16.3%). The number of errors per session ranged from 0 to 13. Colloquial exchanges numbered 951 (48.3%) and formal exchanges 1,016 (51.7%); 482 exchanges (24.5%) contained an idiom, and idiom was more than twice as frequent in colloquial exchanges (34.5%) as in formal ones (15.2%).

Table 1 gives the error rates and odds ratios. Errors occurred in 22.3% of colloquial exchanges and 12.9% of formal ones, and in 31.5% of exchanges containing an idiom against 12.9% of those without. After adjustment, the odds of error were 2.78 times higher when an idiom was present (95% CI 2.19–3.53; p < .001) and 1.58 times higher in colloquial than in formal speech (1.25–2.00; p < .001). Each standard deviation of additional speech rate was associated with 1.21 times the odds (1.07–1.37; p = .003). Each additional ten turns since a glossary term was set was associated with 1.36 times the odds (1.17–1.57; p < .001). Cluster-robust and model-based standard errors differed by between 3% and 6% for every term.

The register estimate fell from an unadjusted 1.94 to an adjusted 1.58 once idiom was in the model, which is what the unequal distribution of idiom between the two registers would lead one to expect. A residual association remained, and its interval excludes 1. Idiom was also implicated more widely than its own error type indicates: 152 of the 343 errors (44.3%) occurred in exchanges containing an idiom, yet only the 56 literal renderings (36.8% of those 152) were of the idiom type. The remainder were omissions, register mismatches or quantity ambiguities occurring in the same turn as an idiom. Because an idiom-type error can occur only in an exchange containing an idiom, part of the idiom association is definitional. In a sensitivity analysis with those 56 errors removed from the outcome (287 errors, 14.6% of exchanges), the error rate was 19.9% (96 of 482) in exchanges with an idiom and 12.9% (191 of 1,485) in those without. The odds ratio for idiom was 1.68 unadjusted (95% CI 1.29–2.20) and 1.48 adjusted (1.14–1.92; p = .003). In the same adjusted model, colloquial register gave 1.57 (1.25–1.98), speech rate 1.26 per standard deviation (1.12–1.42) and turns since a glossary term was set 1.24 per ten turns (1.06–1.45; p = .006).

Turns since a glossary term was set show a pattern that can be read directly from the data. Dividing exchanges into four bands of 0–3, 4–7, 8–12 and 13–37 turns, error rates were 13.4% (74 of 552), 17.0% (87 of 512), 18.0% (77 of 427) and 22.1% (105 of 476). The rise flattened between the second and third bands, so the log-linear term should be taken as a summary of the overall gradient and not as evidence of a uniform slope. For an exchange in formal speech, without an idiom and at the mean speech rate, the model gives a predicted probability of error of 8.2% when a term has just been set and 18.3% when 30 turns have elapsed. For a colloquial exchange containing an idiom, at the mean rate and with a term just set, it gives 28.3%.

4. Discussion

Errors in droid-mediated Huttese-Basic negotiation are frequent in this corpus and are not evenly spread. Roughly one exchange in six contained a material error. Omission, register mismatch and quantity or price ambiguity each accounted for between a quarter and three in ten of the errors, and idiom rendered literally for about one in six. No single failure dominates, which argues against treating droid mistranslation as one problem with one remedy.

Idiom was the strongest predictor in the full model, though part of that strength is definitional. Exchanges containing an idiom had nearly three times the odds of an error, but the literal-rendering type can occur only in such exchanges. With those errors removed the adjusted odds ratio fell to 1.48, still clearly above 1, so the excess was not confined to the literal-rendering type. This matters in practice, because a turn that contains an idiom appears to be a turn in which the interpreter is more likely to lose something else as well. Our data cannot say why. One explanation is that resolving a non-compositional expression consumes effort that would otherwise go to price terms and courtesies, but we did not measure effort, and the same pattern would follow if idiomatic turns were simply denser in content.

Colloquial speech carried higher odds even after idiom was taken into account. Part of the raw difference between registers is a difference in idiom density, and the adjusted estimate shows the remainder. We do not know whether that remainder reflects colloquial vocabulary that the interpreter renders poorly, or a tendency for the register-marking errors defined in our scheme to arise where speakers shift between levels of formality, which earlier work on register marking has emphasised (Tessane, 28 BBY). The corpus can address only the association.

Speech rate was a smaller but reliable predictor, in line with earlier reports of rendering loss at higher rates (Ossel, 25 BBY). An increase of one standard deviation raised the odds by about a fifth, and by about a quarter when idiom-type errors were excluded. We would not extend this beyond the range observed, which was set by the trade-hall speech in the corpus.

The association with turns since a glossary term was set is the finding most relevant to design and to practice. As the count rose, so did the odds of error, and the gradient held after adjustment for register, idiom and speech rate. It fits the operational hypothesis stated in the Introduction, that the interpreter's hold on session-specific terms weakens with intervening turns, and it echoes what bench evaluation has found for context handling (Pellard, 26 BBY). But the corpus does not establish that mechanism. The count is correlated with how far into a session an exchange falls, and later stretches of a negotiation may be more complex, more heated or simply more tiring for the parties, any of which could raise error rates without any change in the interpreter. The flattening of the gradient in the middle bands, seen in the descriptive rates, also warns against reading the estimate as a constant per-turn decay.

Two practical consequences follow for how negotiations might be conducted. If the turns-since-term gradient does reflect a weakening hold on fixed terms, then parties who restate a term at intervals should see fewer price and quantity errors, and this can be tested by assigning refresh intervals to sessions. Flagging idiomatic turns for the interpreter, or asking speakers to paraphrase them, is a second candidate. Neither is recommended here as established practice. Both follow from associations in observational data, and the appropriate next step is a controlled comparison.

5. Limitations

The corpus is a convenience holding. Sessions were recorded where hall operators agreed to release them, and the 19 excluded sessions, particularly the eight with poor audio, may differ systematically from the rest. Because droids were identified by session only, we could not separate the properties of a particular droid from the properties of the hall or the parties, and clustering by session absorbs these without distinguishing them. We do not know whether the results apply to droids of other configurations or to negotiations outside trade halls.

Annotation depends on judgment. Agreement on the presence of an error was good but not perfect, and agreement on type among jointly flagged exchanges was 86%, so some type-level counts carry classification error of a few percentage points. Coding each exchange to a single most consequential type also hides exchanges with several concurrent errors. The idiom predictor is partly tied to the outcome by definition, since idiom-type errors occur only in exchanges containing an idiom, and the sensitivity analysis limits that problem without removing it. The register and idiom codings rest on guidelines and an inventory developed by the authors, and a different inventory would change which exchanges count as idiomatic.

The design is observational, and the four predictors were modelled as linear on the log-odds scale, which the descriptive rates for turns since a glossary term was set do not fully support. We did not model position within the session separately from that count, nor the number of parties, the goods under negotiation or the direction of the bargaining. Only the Huttese-to-Basic direction was annotated. Only the opening stretch of each session was annotated, so the results say nothing about behaviour later in long sessions. Finally, no correction was applied across the four tests, though every interval reported for the adjusted odds ratios excludes 1 by a margin that would survive a conservative correction.

HutteseGalactic Basicprotocol droidsinterpreting errorstrade negotiationcorpus studyOuter Rim

References

  1. Duvane, K. (27 BBY). Omission in consecutive and simultaneous interpreting, a working taxonomy. Coruscant Journal of Interpretation and Translation, 13(2), 88–109.
  2. Orlun, D., & Varsk, P. (29 BBY). Price and quantity terms in spoken Outer Rim contracts. Journal of Huttese and Trade Languages, 7(3), 101–124.
  3. Tessane, I. (28 BBY). Register marking and address forms in Outer Rim trade speech. Journal of Huttese and Trade Languages, 8(1), 3–29.
  4. Ossel, B. (25 BBY). Speech rate and rendering loss in interpreted negotiation. Journal of Huttese and Trade Languages, 10(1), 40–62.
  5. Pellard, R. (26 BBY). Session-context handling in translation droids, a bench evaluation. Coruscant Journal of Droid Engineering, 31(4), 210–232.
  6. Teshnar, L. (30 BBY). Trade halls of the Outer Rim, venues, languages and interpreter use. Outer Rim Planetary Survey Reports, Report 118.
  7. Vess, C. (27 BBY). Fixed expressions in Huttese trade talk, a working inventory. Journal of Huttese and Trade Languages, 9(2), 55–81.
Read this article inside the full journal experience — browse by faculty, search across universes, and explore related work.
Open in Uncited Press →