← State of the Evidence

PPR & Prediction State of the Evidence v0.1

Patient Presentation Rates & Prediction — State of the Evidence

Revised: 2026-08-16 (v2) — supersedes the 2026-08-07 draft. What changed: a correction — Arbon co-authored the Zeitz 2005 comparison, so the two anchor lines of evidence are not independent and are no longer described as if they were; a new section on the denominator problem exposed by the 2026 Purple Guide's capacity definition; the guidance-side convergence on profile-over-headcount; the Meites & Brown threshold evaluation; one new pediatric-event PPR datapoint (Thierbach); and the GP-practice sizing model for multi-day camping events as a candidate additive baseline.

Bottom line for practitioners

Plan against a range, not a point estimate. The best aggregate baseline remains Arbon's 201-event Australian dataset: roughly 1 patient presentation per 1,000 attendees and 0.027 transports per 1,000 across all event types — but individual events deviate from that baseline by an order of magnitude or more in either direction [1,2,3]. If your event recurs, your own historical data will out-predict any published regression model on a day-by-day basis [4]. If it does not recur, use a stratification approach (weather, attendance, alcohol, demographics, crowd intentions) to place the event in a minor/intermediate/major class, and widen your confidence interval sharply for major events, where scoring performs worst [5]. Staff for a mostly-minor, mostly-medical caseload — in one 216-event series, 84% of care was minor/basic and 69% of cases were medical rather than traumatic — while holding contingency for the heat, alcohol, and crowd-behavior modifiers that multiply rates [2]. Remember that on-site physician-level care is itself a lever: at one motorsports venue it cut ambulance transports by 89% [6]. And check your denominator: the population at risk includes the workforce, not just the audience — see below.

What we know

Aggregate rates and regression (Arbon 2001). A 12-month prospective survey of 201 Australian mass gatherings (combined audience >12 million) under a standard reporting format found a patient presentation rate (PPR) of 0.99/1,000 attendees and a transport-to-hospital rate (TTHR) of 0.027/1,000 across all event types. Presentation rates fell slightly as crowd size grew; humidity, crowd mobility, alcohol availability, and bounded venues were associated with higher presentations. Three regression models were derived for prospective prediction [1].

History beats models (Zeitz 2005) — with a provenance note. A prospective head-to-head at a recurrent event compared the Arbon regression model against a retrospective review of seven years of that event's own data. The event's actual PPR was 1.6/1,000 and TTHR 0.07/1,000 — both well above the Arbon all-events averages — and the historical method proved more accurate on a day-by-day basis [4]. Correction carried from the 2026-08-15 metadata audit: the paper's full author list is Zeitz KM, Zeitz CJ, and Arbon P — the comparison's authors include the regression's originator. That is to the finding's credit (the model's own author co-published the result that history beats it), but it means the "Arbon line" and the "Zeitz line" are one research community, not two independent confirmations, and this synthesis no longer describes them as convergent independent evidence. The underlying seven-year series and its follow-up are in the library [7,8].

Stratification works, roughly (Hartman 2009). A retrospective analysis of 55 varied events scored each on weather, attendance, alcohol presence, participant demographic, and crowd intentions. Minor, intermediate, and major events averaged 2.3, 6.3, and 71 patient contacts respectively, with consistent trends for transports. The score correctly predicted resource demand across classes but was least accurate for major events — exactly where the stakes are highest [5].

Event type is a rate multiplier (Milsten 2003). Across 216 events (~9.7 million attendance) in one metropolitan region, overall medical usage was 6.1 per 10,000: baseball 4.85, football 6.75, rock concerts 30 — and a single concert with mosh pits hit 110 per 10,000, a >20-fold spread within one region's data [2].

Scale of the caseload at the extreme. At the largest single-day ticketed concert in North America (>450,000 attendees), 1,870 people sought care — 42 per 10,000 — and no record was kept for 665 of them, illustrating both the load and the documentation problem [9].

A pediatric-event datapoint, recovered. At a fun fair attended by ~100,000 children, overall usage was 19.2 encounters per 10,000 spectators over nine hours, 2% transported, complaint mix dominated by minor trauma (53.6%) with insect bites at 10.4% — the only all-child event PPR we hold, recovered in the 2026-08-15 identifier audit [10].

Conceptually, Arbon's framework — biomedical, environmental, and psychosocial domains combining to produce the PPR, with "latent potential for injury and illness" as a precursor risk state — remains the field's organizing theory [11], extended by the Turris/Lund event-modeling lenses [12].

The denominator problem — new in v2

US presentation rates are almost universally expressed per attendee, where attendance means tickets sold or scanned. The 2026 Purple Guide states, in three separate chapters, that the site population is larger than that: event capacity "will include the number of staff, contractors, guests, performers, volunteers and all those who have a statutory duty to be on site" [13]; pre-site data collection must include "workforce to support the event and breakdown" [14]; and camping chapters provide for crew camping [15].

Three consequences for this domain's core metric:

1. The published rates divide staff-inclusive numerators by staff-exclusive denominators wherever event medical teams treated workers (they generally do), biasing PPR upward by an unknown amount. 2. Build and breakdown are unmeasured exposure windows. The workforce is on site when the audience is not, doing more hazardous work; coverage and rate studies scoped to show days miss it entirely. 3. The workforce risk profile differs — occupational trauma, heat, fatigue, working at height — from the audience profile the models were fitted on.

We know of no study that quantifies any of the three. It is a checkable, publishable methodological point about the field's central metric, sourced to national guidance.

Convergence from the guidance side — and one regulatory evaluation

The empirical literature's conclusion that headcount is a weak predictor now has independent operational support. The Purple Guide argues profile-over-headcount in at least three places: its SAG chapter states that risk "may be greater with events that may not reach" an attendance trigger because "the profile of the audience is as important," and its referral diagram contains no attendance factor at all — ten risk gears (location, traffic, organiser inexperience, criminality, iconicity, profile/behaviour, new event/venue, distance, audience, history) with no headcount gear [16]; its Crowd Management chapter holds that "many crowd behaviours are predictable with an understanding of the crowd's demographics — even matters such as arrival and departure rates" [13]; and its capacity method selects the expected-density figure by risk assessment of the audience profile [14]. Guidance reached by operational experience and regression reached by data agree; US regulation, built almost entirely on headcount triggers, agrees with neither.

And one US threshold has actually been evaluated. San Francisco's 2006 rule required standby ambulances above 15,500 attendees; Meites & Brown measured the result across 47 events — mass-gathering transport at 1 per 59,000 attendees versus a community baseline of 1 per 20,000 residents per six hours (RR 0.15, p<0.001), with 46% of mandated ambulances unused [17]. The one formally evaluated headcount trigger over-provisioned roughly seven-fold; the city's 2025 successor policy nonetheless lowered the trigger six-fold [18]. Threshold-setting and evidence are not currently in contact, and that is a finding for this domain, not only the legal one.

A candidate additive baseline for multi-day events. The Purple Guide's Campsites chapter proposes sizing multi-day camping-event medical resources to the demand on "a GP practice serving a community of similar size" — a residential primary-care model predicting a flat baseline (medication continuity, minor illness, dental, mental health) that acute PPR models structurally omit. For multi-day camping events the two models are plausibly additive; nothing in the literature tests this [15].

What drives the rates

- Weather/heat. Medical usage rates were higher at apparent temperatures ≥80°F (8.1 vs 4.9 per 10,000) [2]; humidity was independently associated with higher presentations in the Arbon dataset [1]. - Alcohol and drugs. Alcohol availability raised presentations in Arbon's regression [1] and is a component of the Hartman score [5]. At sporting mass gatherings specifically, drug and alcohol-related presentations contributed up to 10% of ED visits, with alcohol a factor in up to 25% of ambulance transfers [3]. - Event type. The baseball → football → rock-concert → mosh-pit gradient (4.85 → 6.75 → 30 → 110 per 10,000) is the clearest event-type signal in the literature [2]; concert-genre effects are further explored in unsummarized series [19,20,21]. - Boundedness and mobility. Bounded venues and greater crowd mobility were both associated with higher presentation rates [1]. - Crowd size. Counterintuitively, PPR declined slightly as crowd size increased [1] — a reason to distrust simple linear per-capita extrapolation. - Demographics and crowd behavior. Score components [5] whose independent contributions remain poorly understood; much of the supporting literature is single-event and anecdotal [22]. - The care model itself. On-site physicians reduced transports by 89% (116 → 13; p<0.001), with paramedics able to disposition 52% of patients and protocol-guided RNs a further 39% — the provider mix changes the TTHR you should predict [6]. - Egress is the mirror image. The same audience factors that raise presentations — alcohol, drugs, prams — are named by UK guidance as factors that slow escape below standard egress rates [14]; demand and evacuation capacity degrade together.

What's contested or fragile

Model transportability. The one direct prospective comparison found the Arbon regression outperformed by a simple historical review at a recurrent event [4] — noting again that the comparison shares authorship with the model. Commonly quoted correlation coefficients contrasting the two approaches do not appear in our library's summarized fields and are therefore not reproduced here — a verification target for BA. Later comparative work exists but is unsummarized in the library [23,24,25].

Nonlinearity. Arbon's own group later argued that predicting patient load is an inherently nonlinear problem, that reliable prediction tools are lacking, and that the exact contribution of individual variables remains poorly understood [26] — a candid downgrade of confidence in the 2001 linear regressions by their originator.

Denominators and definitions. Rates are reported in a manner that is "varied, haphazard and author dependent" [27], and inconsistent event and patient characterization makes cross-event comparison difficult [28]. Wide reported ranges — in-event PPRs of 0.18 to 41.9 per 1,000 and transport rates of 0.02 to 19 per 1,000 at sporting events alone [3] — partly reflect true heterogeneity, partly definitional chaos. The workforce-denominator problem above adds a systematic component to that chaos.

Generalizability. The anchor datasets are Australian [1,4,22] or single-region American [2,5]; the field has been repeatedly criticized as dominated by single-event descriptive reports [11,29].

What we don't know

- No meta-analytic PPR estimate exists. Reviews from the 25-year case-report synthesis [30] through systematic reviews of health-service impact [31], ED admissions [32], and stadium/arena presentation factors [33] characterize the literature but none of our holdings yields a pooled, weighted PPR. - Transport-rate determinants are under-studied. TTHR is reported alongside PPR [1,4] and is modifiable by care model [6], but no summarized study models its drivers independently. - A minimum data set has been proposed, not adopted [27,28]; until it is, every new series adds noise as well as signal. Any adopted version should record the workforce population and staff-versus-attendee status of each presentation, or the denominator problem stays unmeasurable. - Crowd demographics, behavior, and many event-type effects remain anecdotal [22]; additional retrospective series exist but are unsummarized in our library [34,35,36,37]. - The WHO research agenda's five public-health directions [38] and Arbon's call for less descriptive, more conceptual research [29] both remain largely unanswered in the prediction space.

How MGMI operationalizes this

MGMI's PPR calculator implements three models, in deliberate tension: arbon_2001_regression (primary-verified; the 201-event derivation [1]), hartman_2009 (stratification scoring [5]), and zeitz_historical (your own event history, preferred where ≥1 prior year exists [4]). Outputs are presented as planning ranges, not point estimates: the documented spread across event types and settings [2,3], the nonlinearity of the problem [26], and Hartman's degraded accuracy at major events [5] make single-number predictions a false precision the literature cannot support. Multi-day camping profiles should surface the GP-practice baseline question as a flagged consideration rather than a computed number — no coefficients exist to compute one [15].

Reading pathway

Start here (in order): 1. [1] — the aggregate baseline and regression. 2. [4] — why event history beats models (shared authorship noted). 3. [5] — practical stratification. 4. [2] — event type, heat, and acuity mix. 5. [27] — why the numbers disagree. 6. [17] — what happened when one threshold was actually evaluated.

Then depth: [11], [6], [9], [26], [22], [12], [28], [29], [31], [33], [3], [38], [10]. Historical/foundational and unsummarized holdings: [39], [35], [30], [40], [19], [41], [7], [8], [34], [24], [25], [21], [23], [37], [36], [20], [32].

Citations

Citation list unchanged from v1 except as noted; additions:

- [4] Zeitz KM, Zeitz CJ, Arbon P. Forecasting Medical Work at Mass-Gathering Events: Predictive Model Versus Retrospective Review. Prehospital and Disaster Medicine, 2005;20(3):164–168. DOI 10.1017/s1049023x00002399. (Full title and third author restored per the 2026-08-15 metadata audit.) - [10] Thierbach AR, Wolcke BB, Piepho T, Maybauer M, Huth R. Medical Support for Children's Mass Gatherings. Prehosp Disaster Med 2003;18(1):14–19. DOI 10.1017/s1049023x00000625. PMID 14694895. - [17] Meites E, Brown JF. Ambulance Need at Mass Gatherings. Prehosp Disaster Med 2010;25(6):511–514. DOI 10.1017/s1049023x00008682. - [18] San Francisco EMS Agency Policy 7010, eff. 2025-04-01 (single-source read; pending visual verification). - [13] · [14] · [15] · [16] Events Industry Forum, The Purple Guide, chapters published 2026-01-26 (chapter-level extractions; coverage in entry notes). - All remaining citations as in v1 (Arbon, Hartman, Milsten, Grange, Feldman, Anikeeva, Ranse, Lund, Turris, Delany, Tam, Michael, DeLorenzo, Richards, Sanders, Baird, Locoh-Donou, Nable, Cannon, Goldberg, Westrol, 2025 reviews).