VULTAX - WHERE UNUSUAL POLYMARKET WINS CONCENTRATE: METHODOLOGY
Companion methods to Vultax Research, version 1.0, 8 October 2026. Adapted from the 7 October aggregate summary.

PURPOSE AND STATUS
Describe where outcomes depart from explicit benchmarks and identify competing explanations. This is exploratory retrospective analysis, not a probability that any person used inside information. The earlier one-page version highlights two results; the attached tables preserve other completed searches, including failures. Source analyses were completed 5-6 October. Differences, percentages and ratios were recomputed from saved outputs on 7 October; this was not a new raw rerun of 714.7 million fills.

DATA AND UNITS
The Qin/Yang, TimeSeventeen/Polymarket-v1 archive covers 1 January 2025 to 28 April 2026: 714,691,643 fills, 834,732 markets, 2,404,822 wallet addresses. Fills, wallet addresses and bets are not unique people or independent events. Polymarket metadata and observed payouts supplement the archive. Removing sports, esports and fixed-clock crypto leaves 135,284 markets, 122,509 resolved within the tape and 118,158,616 fills. Different analyses use different denominators; the original tight episode screen includes excluded sports strata. Current Gamma metadata are not historically point-in-time tags.

A. EARNINGS / AUDITOR COMPARISON
Main roster: 891 markets, 756 resolved, 423 companies and 756 company-reporting cycles. Annual audit firms come from PCAOB AuditorSearch/Form AP; company/ticker identity was resolved using SEC public entity information and exact issuer-name matches. Annual-auditor mapping does not identify quarterly-review staff or prove personal access.

The known group comprises thirteen published addresses (initial nine plus supplementary four). It is not a reconstruction of every reported participant. Pool the earliest directional entry across group members for each company/reporting cycle, once per cycle, rather than summing duplicated wallet bets. Count binary wins of that initial direction, not profitable markets or realised trading P&L. Separate KPMG clients from clients of other known auditors.

For N counted positions with recorded entry probabilities p_i and binary outcomes y_i:
 E = sum(p_i), W = sum(y_i), excess wins = W - E.
 Expected rate = E/N; observed rate = W/N.
 Percentage-point gap = 100*(W-E)/N.
 Observed/expected multiple = W/E, where E > 0.
These comparisons are not stake-weighted returns. Entry prices are used as calibrated-probability benchmarks, not assumed 50:50 probabilities. Adding expectations does not require independent outcomes; probability tails and uncertainty do require assumptions about dependence and sample selection.

Results: KPMG clients W=32, N=37, E=23.331: 8.669 extra wins, 86.4865% observed versus 63.0568% expected, +23.4297 percentage points, 1.37157 times expectation. Other known auditors W=28, N=55, E=30.704: -2.704 wins, 50.9091% versus 55.8255%, -4.9164 points. The E totals are published to three decimals in the source report. KPMG has 74 of 423 mapped companies, while 80.55% of group main-earnings directional cost is on its clients. Company share is context, not a uniform random-choice model.

Sensitivity: 12 cuts per auditor = tail ceilings 0.01/0.001 x minimum inside cycles 3/5/8 x cost share 50%/80%. Compare selection counts with 1,000 company-label permutations, keeping all a company's quarters together and preserving firm company counts. KPMG exceeds its own 95th percentile in five cells, other auditors in none. The cells overlap; no multiplicity-adjusted claim follows. Shuffles do not preserve sector or company size. A separate earnings-call mentions roster has 906 markets, 885 resolved and 56 companies; it produced no KPMG flags under the corresponding search.

B. DOMINANT-SUBJECT POPULATION REFERENCE
Use the retained population above. Union Gamma market and event tags. Trim and lowercase slugs; remove empty tags, broad tags (sports, esports, politics, crypto, finance, business, economics, economy, science, tech, culture, pop-culture, world, news, elections, global-elections, us-politics, geopolitics, other), and whole-token administrative tags hide/featured/trending/new/recurring/all. Other retained tags can still be imperfect semantic descriptions.

For each wallet, choose the retained tag with the largest lifetime directional acquisition cost share. Count each market once within a tag; break ties by distinct event count then lexical tag. The denominator includes ALL observed lifetime directional costs, including excluded categories and unresolved positions. This is directional turnover, not a deposit balance or maximum capital at risk.

The displayed reference is primary cell 28: at least 80% lifetime cost on the dominant subject; first subject entry within three days of first observed tape activity; first observed activity on/after 1 February 2025 to reduce left-censoring; at least three binary resolved subject events; at least $2,500 resolved subject cost. Account age means observed history, not legal account creation. For each event use the earliest binary resolved market entry, breaking timestamp ties by stored market identifier. Fractional payouts can contribute cost/P&L but not binary trial counts.

Eligibility precedes the success-tail test. For a wallet's event entries compute the Poisson-binomial upper tail P(sum independent Bernoulli(p_i) >= observed wins); keep tail <=0.01. This is a model-based screen criterion, not a 99% probability of misconduct. The reference has 22,322 eligible wallets and selects 321.

SIMULATED COUNT OF QUALIFIERS
500 simulations, seed 20261005. Hold observed entries, entry prices, dominant subjects and non-tail eligibility fixed. Draw one uniform U per Gamma event per run and reuse it across wallets and cells: reference-direction wins if U<p; opposite-direction wins if U>=1-p. Distinct Gamma events have independent draws. Reapply the same success-tail threshold to each simulated wallet and count qualifiers. Expected count is the simulation mean; displayed band uses linearly interpolated 5th and 95th percentiles.

Result: simulated mean 51.252, 5th-95th percentiles 27-87; observed 321. Excess 269.748; observed/expected 6.26317; observed eligibility rate 1.43804% versus simulated mean rate 0.229603%. These are WALLET selections, not wins. Dominant weather and daily-temperature groups account for 162 of the 321 selections; this is a descriptive domain split, not a finding that every such trader used a public model.

The shared-uniform construction preserves fair-price marginal probabilities and some within-event crowd dependence. At heterogeneous entry prices, it does not recreate one coherent literal binary settlement for every entry. Cross-event correlation, adaptive trading and miscalibrated prices are not simulated. This model mismatch can produce excess selections. The simulation is neither a calibrated false-positive rate nor a count of insiders.

SEARCH BREADTH AND NEGATIVE RESULTS
Eight initial topic/time rules cross market/event/event-or-series/shared-tag grouping with 48/72 hours. The inherited large-low-price screen uses a rolling-hour stake at least $2,500, prices 5-35 cents, observed tape age under 14 days, first observation on/after 1 February, at most one unrelated prior market, and at least 90% directional-cost concentration. Its retrospective outcome clock starts the final price run staying at or above 95 cents, not formal settlement. The corrected strict 48-hour screen has 1,316 bets, 264 wins versus 297.3625 expected; matched gap -5.40 cents per share, interval -10.53 to -0.07. Broader versions recover studied military examples without a clearly positive matched average gap. Established controls require at least 90 observed days and ten prior markets; matching uses category, price band, half-year and clock bins. Intervals use 2,000 market bootstrap draws, leaving possible dependence between related markets.

The first population grid has 216 combinations: event/series/narrowest retained tag x 80/90/95% concentration x 2/3/5/10 resolved events x tail ceilings 0.01/0.001/0.0001 x resolved cost $0/$2,500. None retained the studied named case wallets.

The improved overlapping-subject grid has 72 primary cells: 50/80/95% share x within-three-days/no age gate x 3/5 events x tail ceilings 0.05/0.01/0.001 x $0/$2,500. Another 144 censor-stratum rows assess the same cuts in history-coverage groups, not independent experiments. The displayed 80% reference was specified in the exploratory protocol; it is not an independently held-out test. No single cell retained all studied case wallets.

Other source studies examined information-holder labels, public-news timing and funding/owner links. Their aggregate tables remain attached. The early-holder/no-public-model subset has 378 bets, 155 wins versus 73.706 expected, but the separate matched gap is +7.6 cents with interval -6.1 to +16.8 and only 79.1% matched. Labels have incomplete validation and coverage. The military timing audit covers 224 flagged bets, 189 wallets and 48 markets: 132 of 133 wins had a located hard public signal within 72 hours, one uncertain. Retrospective source timing is imperfect and does not establish a trader's actual information. Funding checks did not reconstruct a link among the thirteen published KPMG accounts. These qualifications prevent interpreting the headline departures as a validated causal explanation.

PROVENANCE AND REPRODUCTION
This public archive contains aggregate tables for the publication. headline-data.csv contains three derived comparisons with explicit units: two company-cycle win rows and the reference wallet-selection row. strict-screen-comparisons.csv contains eight corrected topic/time comparisons. The sensitivity and subject tables preserve the larger aggregate searches, with named-account membership fields omitted. Recompute each headline difference, rate and multiple with reproduce-headlines.py and the formula above. The manifest fingerprints each public file and selected original source outputs. The manifest identifies retained source reports by hash. The public package does not contain those private working reports or a full raw-data environment. This package is an auditable aggregate summary, not a complete raw-data reproduction environment.

PUBLIC SOURCES
https://huggingface.co/datasets/TimeSeventeen/Polymarket-v1
https://pcaobus.org/resources/auditorsearch
