Supported finding
238 of 4,710 feature markets appear anywhere in the trade file, a 5.05% intersection. The extracts cover unequal periods, so this does not establish contemporaneous quote/trade coverage.
05 / Market microstructure
A corrected descriptive analysis of preserved market-minute features and trade records. Inspect activity, source-field spreads and the limits of dataset overlap.
Question
How much of the feature-market population appears in the supplied trade file?
Fixed comparison: Feature extract, 6–11 March 2026 UTC; trade extract, 17 July 2025–27 March 2026 UTC. Controls below do not change this summary.
238 of 4,710 feature markets appear anywhere in the trade file, a 5.05% intersection. The extracts cover unequal periods, so this does not establish contemporaneous quote/trade coverage.
Before comparing source populations, an analyst should retain the intersection denominator and each file’s time window. Restrict this edition to descriptive source-field summaries; execution-cost comparisons need stronger identity and timing evidence.
Sampling and upstream field construction are unknown. A market ID match does not align a token, execution or pre-trade quote. No market-wide, effective-spread, causal-impact or profitability conclusion follows.
Feature bars cover 6–11 March 2026 UTC. Controls summarize only recorded bars; missing market-minutes are not filled with zeros.
Historical assignment extracts with unknown sampling and upstream feature definitions. These summaries do not estimate all Polymarket activity, execution cost, price impact or profitable trading opportunities.
Start here Switch Measure to compare activity, source-field spread or volume. Then inspect the exact table.
Loading the corrected, validated extract…
Full extract totals. Grouping and measure controls affect the chart and its table below.
Each mean weights observed bars equally. Markets with more recorded bars contribute more. The hourly view combines the full six-day feature extract; it does not establish a recurring schedule.
Only the intersection belongs in a coverage numerator. Dataset periods and populations remain distinct.
All 123,895 metadata category strings are blank in the source. Matching metadata cannot recover a classification that was never supplied. The earlier diagnosis of categories being lost in a join was not supported.
No substitute category labels are generated.
Trades span 17 July 2025 to 27 March 2026 UTC. The file contains 4,126,076 records across 23,016 markets, including 177,477 exact repeated rows.
No execution ID distinguishes repeated ingestion from separate executions. Recorded-row aggregates retain these rows. BUY includes both Yes and No assets and does not measure bullish sentiment.
The 238 feature markets appearing anywhere in the trade file are 5.0531% of 4,710 feature markets. This is an identity intersection across unequal periods, not a contemporaneous execution-to-quote match. Metadata’s as-of date is unknown.
The corrected edition uses every row of the three checksum-matched inputs. The original analysis is retained separately for provenance.
The original percentage logic used unrelated metadata or trade populations. The new coverage table shows both numerator and denominator, checks unique keys and keeps metadata availability separate from category availability.
The original interquartile-range cap had an upper bound of zero in these sparse inputs, erasing positive activity. The new edition does not cap volume. It checks all-zero cases and retains tied activity extrema.
At an absolute tolerance of 0.0001, 1,135 feature bars differ from total volume = buy volume + sell volume, and 684 differ from order-flow imbalance = buy volume − sell volume. No value is altered to force either identity.
These are numerical diagnostics of supplied fields. Unknown upstream construction prevents a universal economic interpretation.
Feature bars lack token IDs and pre-trade quote timestamps. Trades cover both Yes and No assets. A market/minute join therefore does not support effective spread or causal price impact. Targets lack construction and horizon documentation.
Category comparisons, target prediction, resolution relationships, distributional classification, dollar notional, strategy returns and market-wide inference are excluded.
The download includes code, pinned package versions, regression tests, validation, source checksums and complete aggregate outputs. Raw inputs and market identifiers are excluded.
R code and aggregates included. Full replay needs authorized exact source files.
R and the dependencies listed in the package.
Authorized exact Parquet sources required; raw records are not included.
Read the baseline tables immediately, or restore authorized source copies and replay the corrected analysis.
Download reproduction ZIPZIP · 549 KBFollow the README to restore dependencies and the three exact Parquet inputs, then run:
Rscript --vanilla test.R Rscript --vanilla run.R INPUT OUTPUT Rscript --vanilla validate.R INPUT OUTPUT
The original extracts have no supplied acquisition URL or redistribution licence. The package records filenames, sizes and hashes; it does not invent a download source.
Aggregate results are inspectable from the package. A complete rerun requires access to the governed source files.