← back to dashboard

Scan2029 — Methodology & Reasoning

How the Indonesia 2029 presidential tracker collects data, forms estimates, and where human judgment enters. Last revised 10 August 2026.

1. What this system claims — and doesn't

Scan2029 estimates, for any given day, the probability that each tracked figure wins the 2029 Indonesian presidential election. The estimate is not a prediction of a fixed future; it is a snapshot of evidence (polls, news coverage) passed through an explicit model of how Indonesian elections actually get decided — nominations and coalition math first, popularity second.

Until coalitions crystallize (~2028), the honest reading of any number here is a trend with wide error bars. Indonesian history is unkind to early forecasts: Jokowi was a provincial mayor three years before winning the presidency, and Dedi Mulyadi was absent from national polling eighteen months before topping a July 2026 survey. The model encodes this humility directly — see the noise term (§5) and the "field" entrant (§6).

2. Tracked candidates

The roster is the top four by national polling plus two structural candidates who hold permanent seats despite little or no national polling, because their nomination path is credible: Pramono Anung (Governor of Jakarta, acceptable to both the Megawati and Jokowi wings of PDI-P) and Sherly Tjoanda (Governor of North Maluku, elected 2024 on an eight-party NasDem-led coalition, unaffiliated and courted by several parties — the type of outsider provincial executive Indonesian politics has repeatedly elevated). Ganjar Pranowo was reviewed and excluded. Everyone else — AHY, Ganjar, ministers, unknown 2028 entrants — competes in the simulation as an aggregate "field" entrant (§6), and a tripwire (§3) watches for breakout names that should be promoted onto the board.

3. Data pipeline

LayerSourceCadence
PollsNational electability surveys (Indikator, SMRC, Litbang Kompas, Poltracking, LSI, Median, Indekstat, IPI, IPO, Index Politica), entered from public releases and a weekly automated sweep of Wikipedia's 2029 polling pageWeekly + manual
NewsGoogle News (Indonesian editions) queries per candidate, ~50–150 articles/day from detik, Kompas, Tempo, CNN Indonesia, Antara, and othersDaily 05:00 WIB
LLM taggingEach article is machine-read (GPT-4o-mini) into structured tags: per-candidate sentiment (−1…+1), themes, voter segment implicated, stance of cited voices, and event type (coalition move, scandal, ruling, endorsement)Daily
SynthesisA larger model (GPT-4o) writes each candidate's plain-English state of play, strengths, and weaknesses from the last 30 days of tagsWeekly
TripwireCounts of untracked names appearing in presidential context — the early-warning system for the next KDM-style breakoutDaily
LLM-extracted tabular data is treated as untrusted until verified. An early version of the poll sweep corrupted column alignment (assigning one candidate's percentage to another); the pipeline now parses tables structurally, normalizes pollster names, and deduplicates by pollster + fieldwork month. See §8 for the full audit history.

4. The poll aggregate

Each candidate's headline "poll aggregate" is a weighted average of all national capres (first-choice) survey results:

weight = pollster_quality × 0.5^(age_days / 75)
Known limitation: top-of-mind (open recall) and closed-list (simulated ballot) surveys are currently aggregated as one series, although they measure different things — open recall structurally favors incumbents. A methodology tag per poll is the planned refinement.

5. Monte Carlo simulation

The score from §4 measures popularity today. Winning in 2029 requires surviving nomination politics and thirty more months of events. The model handles this with a 10,000-run Monte Carlo:

  1. Each run samples every scenario (§6) true/false by its probability, multiplying candidate scores by the scenario's effects.
  2. Each candidate's score is then perturbed by lognormal noise (σ = 0.55) — wide enough that a candidate at 8% today still occasionally reaches the mid-20s in simulation. This is the mathematical expression of "2.5 years is a long time": it reflects how far early polls historically sit from final outcomes.
  3. The winner of each run is the highest adjusted score; win probability is the share of runs won. The p10–p90 band shown on each card is the spread of simulated vote shares.

6. Scenarios — where judgment enters, explicitly

This is the layer most forecasts hide and Indonesia most requires. Raw polling answers "who is popular"; scenarios answer "who is even on the ballot, with what machinery." Every scenario, its probability, and its effects are published on the dashboard and editable — they are analyst judgment calls, stated in the open so you can disagree with a specific number rather than a black box.

Current reasoning behind the key entries:

Scenarios support dependencies: a scenario can require another (consolidation behind Anies cannot fire without his ticket) or be gated by one (Dedi's vehicle problem only exists if Prabowo runs). Impossible combinations are excluded from every simulation.

The "field" entrant carries a fixed score of 12 — roughly the current combined polling of AHY, Ganjar, and rising ministers — and competes in every run. Its win share (~8–10%) is the model saying: this far out, the eventual winner may not be on the board yet.

7. What would change the numbers most

8. Audit history — errors found and fixed

Forecast credibility comes from showing corrections, not hiding them.

DateIssueResolution
9 Aug 2026LLM poll sweep corrupted column alignment (e.g. Anies stored at 1.3% vs actual 8.2%; SMRC July leaders flipped), producing an indefensible 0% for AniesPoll table rebuilt from source-verified rows; extractor rewritten to parse table structure; caught after a user challenge
9 Aug 2026Single-poll dominance: one fresh survey drove an 80% win probabilitySparse-data momentum damper; longer half-life; wider simulation noise
10 Aug 2026Independent adversarial review (separate AI agent): scenario table structurally biased (Dedi had no downside branch), three missing Dedi poll rows, newest IPO poll absent, incoherent scenario combinations, forced-choice bias across only 5 candidatesMissing rows + IPO added; Dedi nomination-vehicle scenario; scenario dependencies; "field" entrant; sweep hardening
21 Aug 2026Sherly Tjoanda (Governor of North Maluku) absent from the board despite a structural nomination case comparable to Pramono'sAdded as a sixth tracked figure with a permanent seat; no national capres polling exists for her yet, so she enters at the no-poll floor and is carried by news coverage and the model's noise term
10 Aug 2026Prabowo's biography (five pursuits, patrimonial party control) underweighted at P(runs)=0.65Raised to 0.75; Dedi's vehicle scenario gated on Prabowo running and cut to 35% — flipped the leader from Dedi 43% to Prabowo ~40%

Known open limitations: Pramono's aggregate rests on one minor-pollster reading while other surveys place him near zero (his probability is likely optimistic); LLM sentiment carries positivity bias (mitigated by the ±8% cap); poll question methodologies are mixed (§4).

9. Stack

Cloudflare Worker + D1 (SQLite), scheduled crons (daily 05:00 WIB pipeline; Monday poll sweep and syntheses), OpenAI GPT-4o-mini/GPT-4o for reading Indonesian media at scale, and a dependency-free hand-rolled Monte Carlo. All model math is deterministic, versioned code; all judgment (scenario probabilities, pollster weights) is data, published on the dashboard and editable without redeploying.