More Informants, Less Information: How the Accuracy of a Broker's Assessment Depends on the Number and Independence of Informants in Bothawui Brokerage Ledgers, 30–26 BBY
Abstract
Brokers who consult several informants assume that agreement among them is evidence. That assumption holds only if the informants learned what they know from different places. We asked how the accuracy of a broker's assessment depends on the number of informants consulted and on how independent their sources were. We drew on the assessment registers of nine brokerage houses in the Bothawui Brokerage Ledgers, covering forecasts of named trade and political events issued between 30 and 26 BBY and resolved by 25 BBY. Of 471 entries screened, 424 met the inclusion criteria. Each was coded for the number of informants and for the fraction whose accounts traced to a shared upstream origin; coding agreement was high (kappa 0.88). In a logistic model, each additional informant multiplied the odds of a correct call by 1.77 (95% CI 1.45–2.17) when all sources were distinct, but by only 1.25 (1.08–1.45) when the number of distinct origins was half the number of informants; the ratio of these per-informant odds ratios was 0.71 (0.57–0.88; interaction p = .002). The finding held under house fixed effects and when single-informant assessments were excluded. Stated confidence rose with the number of informants and did not fall with the shared-origin fraction, and the mean Brier score was 0.216 (95% CI 0.183–0.246). Larger panels were associated with more accurate calls mainly when their sources were distinct.
1. Introduction
Bothans are renowned across the galaxy as gatherers of information, and Bothawui, their homeworld, is where a good deal of that reputation was earned. A reputation, however, is a statement about outcomes. It says nothing about the mechanism by which a broker who has collected many accounts turns them into one reliable assessment, or about how many accounts are enough. This paper examines that mechanism in a limited setting: brokers who forecast the outcome of a named trade or political event, and who report how many informants they consulted. We use the term brokerage house throughout for the firms and family concerns that keep assessment registers in the Bothawui ledgers. The term is ours, and we make no claim about how such houses are organised beyond what their registers record.
Consulting more informants is the ordinary remedy for uncertainty, and there is good reason to expect it to work. If informants err independently, the errors partly cancel and the aggregate is better than any single account. Studies of expert panels in commodity markets report that forecasts improve with panel size at a declining rate (Haleth, 28 BBY). The argument fails when informants are not independent. Two informants who received their account from the same upstream observer, or who read the same document, are one source counted twice, and their agreement adds little. Work on correlated error shows that the effective size of a panel of assessors can be far smaller than its head count (Pell-Aroun, 27 BBY), and studies of relayed reports document how often accounts in a chain descend from a single origin (Sarnen, 29 BBY). What has not been shown is how much this costs a broker in accuracy, or whether brokers adjust their stated confidence when they should.
Our vantage is 24 BBY. The records analysed end in 25 BBY, and all analysis was completed in the following year. We pose three questions. First, does the accuracy of an assessment rise with the number of informants consulted? Second, and centrally, is the gain from each added informant smaller when the informants share upstream origins? This is the prediction of an effective-sample account, and it corresponds to a negative interaction between the number of informants and the shared-origin fraction. Third, as a secondary question, does the broker's stated confidence track informant count, independence, or both, and how well calibrated is it as a probability? We report every analysis that we ran.
2. Methods
Records. The source was the set of assessment registers held in the Bothawui Brokerage Ledgers for nine brokerage houses (Bothawui Brokerage Houses, 25 BBY). The houses were those whose registers recorded, for each assessment, the number of informants consulted, the stated confidence, and a later entry resolving the outcome. Assessments were issued between 30 and 26 BBY, and outcomes were taken from resolving entries made no later than 25 BBY. An assessment was a forecast that a named trade or political event would or would not occur within a stated period, typically between 30 and 180 standard days. Typical events were the revision of a named tariff, the seating of a named delegation, and the settlement of a named route dispute. The choice of event types follows earlier work on tariff revision as a forecastable event (Ondrel, 26 BBY).
Sample. We screened 471 register entries. We excluded 47: 29 whose outcome had not been resolved by 25 BBY, and 18 for which the number of informants was not recorded. That left 424 assessments for analysis, contributed by houses of between 26 and 70 entries each.
Coding of independence. For every informant in an assessment, the register usually carries a short note of where the informant's account came from. Two coders followed the Institute's provenance manual (Bothan Institute of Information Sciences, 26 BBY) and traced each account to its upstream origin, meaning the first-hand observer or original document from which it descended. Two informants shared an origin if their accounts traced to the same one. For an assessment with k informants and m distinct origins, we call m the number of effective independent sources and define the shared-origin fraction as S = 1 − m/k. Thus S = 0 when every account has its own origin, and S approaches 1 as more informants collapse onto a single origin. An account whose origin could not be traced was counted as its own origin, which biases S downwards. Coders saw the register note but not the outcome field. A random subset of 120 assessments with two or more informants was coded by both coders, comprising 1,268 pairwise judgments of whether two informants shared an origin. Agreement was measured by a chance-corrected coefficient (kappa; Cassel, 30 BBY). Resolutions were read by one coder, and a second reader checked a random 60 of them.
Outcome and stated confidence. An assessment was scored correct if the side the house favoured was the side that occurred. The stated confidence, c, was the probability the house recorded for its favoured side, which lies between 0.50 and 0.97 in these registers. Because c is a probability for the favoured side, the Brier score, the mean squared difference between c and the 0/1 outcome, is a proper scoring rule for it (Marvane, 29 BBY). A house that always stated 0.50 would score 0.25, and lower scores are better.
Models. Correctness was modelled by logistic regression on the number of informants (centred at four, so that the intercept describes a typical assessment), the shared-origin fraction, and their product. The interaction coefficient is the log of the ratio by which the per-informant odds ratio changes for a unit increase in S. It is a ratio of odds ratios, not a ratio of slopes on the probability scale. We report Wald intervals, a likelihood-ratio test for the interaction, and Akaike's criterion for comparing this model with three alternatives: informants alone, informants plus S without the interaction, and a model in m alone on the log scale. Intervals for contrasts were computed from the full covariance matrix of the fitted coefficients.
Sensitivity analyses. Assessments cluster within houses, whose practices differ. We therefore refitted the model with house fixed effects, and with cluster-robust errors on the nine houses, which is too few clusters for those errors to be trusted and is shown for completeness. A single-informant assessment has S = 0 by construction, so the shared-origin term is identified only from assessments with two or more informants, and a stratum in which no sharing can occur may distort the interaction. We refitted the model on the 387 assessments with at least two informants. For calibration, confidence intervals for the Brier score and for the gap between mean stated confidence and mean accuracy were obtained by resampling houses, with several thousand replicates; intervals within strata (Tables 1 and 3) resample assessments, and Wilson intervals are used for proportions. Because stated confidence and independence both depend on informant count, we also regressed correctness on the log-odds of c and on S together. All figures below were computed by script from unrounded values.
3. Results
Of the 424 assessments, 290 were correct (68.4%; Wilson 95% CI 63.8–72.6). The number of informants ranged from 1 to 8 (mean 4.10, SD 2.06), and 37 assessments rested on a single informant. Among the 387 with two or more, the mean shared-origin fraction was 0.33 (SD 0.25, maximum 0.88), and 96 (24.8%) had entirely distinct origins. The number of effective independent sources averaged 2.70 (SD 1.59; median 2; range 1–8). Accuracy varied across houses, from 51% to 89%. Coders agreed on 94.2% of the 1,268 pairwise origin judgments (kappa 0.88), and the two readers agreed on the resolution of 57 of 60 entries (kappa 0.90); the three disagreements were settled by a third reading.
Accuracy rose with panel size, and the rise was steeper where sources were distinct (Table 1). With one or two informants, the 110 assessments were correct 46.4% of the time (95% CI 37.3–55.6), which is not distinguishable from chance. With six to eight informants the corresponding figure was 83.6% (75.6–89.4, 92 of 110). Within the six-to-eight band, assessments with a shared-origin fraction of 0.25 or less were correct 90.2% of the time, against 79.7% for those above 0.25. At two or three informants the two strata were indistinguishable, at 58.8% and 58.1%.
The regression supports the effective-sample account (Table 2). At S = 0, each additional informant multiplied the odds of a correct call by 1.77 (95% CI 1.45–2.17, p < .001). The interaction was negative: the per-informant odds ratio was multiplied by 0.50 (0.32–0.78) for each unit increase in S, and by 0.84 (0.75–0.94) for each 0.25 increase (likelihood-ratio χ2(1) = 9.39, p = .002). The per-informant odds ratio therefore fell to 1.49 (1.30–1.71) at S = 0.25 and to 1.25 (1.08–1.45) at S = 0.50. The ratio of the odds ratio at S = 0.50 to that at S = 0 was 0.71 (0.57–0.88).
Taken alone, the shared-origin term had little effect in a small panel. For a rise of 0.50 in S, the odds of a correct call at two informants were 1.13 times as large (0.65–1.98, p = .66), 0.57 times as large at four informants (0.36–0.90, p = .016), and 0.14 times as large at eight (0.05–0.42, p < .001). In probability terms the fitted accuracy at eight informants was 97.4% for fully distinct sources and 83.8% at S = 0.50, and at six informants it was 92.1% and 76.8%. At four informants it was 78.9% against 67.9%. Adding S to a model of informant count without the interaction improved the fit only weakly (likelihood-ratio χ2(1) = 3.24, p = .072), so the effect of independence is visible chiefly through how it changes the value of further informants. The full model had the lowest Akaike criterion (492.7), against 501.3 for informant count alone, 500.1 for count plus S, and 498.4 for a model in log m alone. In that last model, observed accuracy rose from 53.0% with one effective source (61 of 115) to 92.6% with six or more (25 of 27), and each doubling of effective independent sources multiplied the odds by 2.09 (1.62–2.71). When k was added to log m it kept a modest effect of its own (odds ratio 1.18 per informant, 1.01–1.37, p = .034; criterion 495.7), so the number of effective sources does not replace informant count entirely.
Every check preserved the interaction. With house fixed effects the interaction odds ratio was 0.50 (0.32–0.80; likelihood-ratio χ2(1) = 8.60, p = .003), and the per-informant odds ratio at S = 0.50 was 1.24 (1.07–1.43). With cluster-robust errors on the nine houses it was 0.50 (0.34–0.72; p < .001), and that interval should be read with the caution given above. Restricted to the 387 assessments with two or more informants, where the shared-origin term is not tied to the single-informant stratum, the interaction was stronger, 0.40 (0.23–0.68; likelihood-ratio χ2(1) = 11.67, p < .001).
Stated confidence was clearly related to the number of informants (r = 0.74) and rose about 4.1 percentage points with each one. Among assessments with two or more informants, its correlation with S was slightly positive (r = 0.15, p = .003; not adjusted for informant count), so houses did not lower their confidence when sources overlapped. Overall the mean Brier score was 0.216 (95% CI 0.183–0.246 resampling houses), which is 13.8% below the score of 0.25 for a constant statement of 0.50. Mean stated confidence was 76.2%, against an accuracy of 68.4%, an excess of 7.8 percentage points (95% CI 0.6–14.2). When correctness was regressed on the log-odds of the stated confidence, the coefficient was 0.53 (0.23–0.82), below the value of 1 expected of a calibrated statement (p = .002), so the spread of stated confidence was wider than the spread of accuracy. The shared-origin fraction had no detectable further effect once confidence was in the model (odds ratio 0.74, 0.33–1.71).
Table 3 breaks calibration down by independence. Among assessments with two or more informants, the confidence gap and Brier score were larger where S was larger, from 4.0 points and 0.197 in the least-shared stratum to 11.7 points and 0.255 in the most-shared. These differences are suggestive only. The difference in Brier score between the most-shared and least-shared strata was 0.058, with a house-resampled interval (−0.032 to 0.121) that includes zero. The single-informant stratum, with S = 0, was worse than any of them (a gap of 13.2 points, Brier score 0.266, with wide intervals), so the pattern is not a clean function of independence.
4. Discussion
In these registers, the value of an additional informant depended on where the informant's account came from. When every informant in a panel traced to a distinct origin, each added informant increased the odds of a correct call by about three quarters. When the number of distinct origins was half the number of informants, the increase fell to about a quarter. That fits an effective-sample account: an informant whose account duplicates one already held contributes little, so the count that matters to accuracy is nearer to the number of distinct origins than to the number of people consulted. The interaction held under house fixed effects and when single-informant assessments were excluded, and its interval under cluster-robust errors is shown but not relied on.
The evidence is stronger for larger panels than for small ones. At two informants, shared origin made no detectable difference to fitted accuracy, and at two or three informants the observed accuracies in the two strata were almost the same. A penalty for shared origin appears when the panel is large enough for overlaps to be numerous. We do not treat it as a demonstrated threshold, because the interaction was estimated as a single linear term and other forms were not tested.
A count of effective sources did not entirely replace the head count. Informant count retained a modest effect of its own when log m was in the model. Coding of origin was imperfect and, where an origin could not be traced, biased towards independence, so some of what looks like an effect of informant count may be undetected sharing. The data cannot separate this from differences in how well placed informants are.
Calibration is the secondary finding, and the evidence there is thinner. Houses stated higher confidence when they consulted more informants, and the spread of their stated confidence was wider than the spread of their accuracy. Their confidence did not fall with the shared-origin fraction, and if anything it rose slightly. Among panels of two or more informants, overconfidence was greatest where sharing was greatest, but the difference in Brier score between the extreme strata has an interval that includes zero, and the single-informant stratum was worse than any of them. What the data support is narrower: stated confidence responded to the number of informants consulted, and the registers show no lowering of it where sources overlapped (an unadjusted correlation of 0.15).
For practice, the finding argues for spending effort on the origin of accounts, not only on their number. On these estimates, an informant from a wholly new origin adds more than one from an origin already represented. Our data are observational, and they identify an association between panel structure and accuracy in past assessments, not the effect of changing the panel for a given question.
5. Limitations
The data are observational and come from nine houses, all of which kept registers detailed enough to be included. Houses that record where their informants obtained their accounts may be more careful than the rest, and their accuracy and calibration may not describe brokerage generally. The number of informants was chosen by the brokers, and a broker may consult more informants on questions that appear harder. That would tend to weaken the apparent benefit of larger panels rather than to strengthen it, but it means that the model does not isolate the effect of panel size. The base rates of the events forecast were not consistently recorded, so we cannot say how much of the accuracy above 50% reflects skill and how much reflects the predictability of the questions asked.
Coding of independence rests on the notes in the registers. Where a note was absent or vague, the origin was treated as distinct, and this understates sharing. Our coders agreed well with each other, but agreement between two coders using the same notes does not show that the notes are complete. The shared-origin fraction and the informant count are also related by construction, since S can only be nonzero when there are at least two informants and can only be large when there are many. We addressed this by restricting to assessments with two or more informants, but the two quantities cannot be made fully separate.
Statistical precision is limited. The cluster-robust intervals rest on nine clusters, and the fixed-effects analysis treats house differences as constant. The pairwise origin judgements used for the agreement figure are clustered within assessments, so the kappa interval, if computed, would be wider than a count of pairs implies. The one-informant and two-informant strata are small and near chance, and the stratum-level calibration figures have wide intervals. Finally, the events forecast were those that houses chose to record, and whether the pattern holds for other subjects is unknown.
References
- Bothawui Brokerage Houses (25 BBY). Assessment registers of nine brokerage houses, with informant counts, stated confidence and resolving entries. Bothawui Brokerage Ledgers, Registers H1–H9, entries issued 30–26 BBY.
- Sarnen, L. (29 BBY). Relay chains in reported information, a taxonomy of upstream sharing. Bothan Institute Transactions on Information Networks, 3(1), 22–47.
- Bothan Institute of Information Sciences (26 BBY). Provenance coding manual for brokerage registers. Bothan Institute Transactions on Information Networks, Technical Note 4.
- Pell-Aroun, T. (27 BBY). Correlated error and the effective size of a panel of assessors. Proceedings of Applied Speculative Statistics, 10(2), 101–128.
- Marvane, E. (29 BBY). Quadratic scoring of probabilistic forecasts and its decomposition. Proceedings of Applied Speculative Statistics, 8(1), 3–19.
- Cassel, J. (30 BBY). Chance-corrected agreement for paired judgements. Proceedings of Applied Speculative Statistics, 7(2), 55–70.
- Haleth, R. (28 BBY). Expert panels and forecast accuracy in commodity markets. Muunilinst Economic Analysis Bulletin, 15(2), 88–113.
- Ondrel, K. (26 BBY). Tariff revision as a forecastable event, evidence from Senate schedules. Journal of Galactic Public Finance, 13(3), 201–226.
Open in Uncited Press →