Why price assumptions in backtests matter: definition and context
Backtests often start with a set of simplifying price assumptions: midpoints, zero commissions, and instantaneous, full-size fills. These idealizations lower implementation complexity but create a systematic gap between simulated and realizable performance when spreads, queueing and venue effects matter, so realistic backtests need to anchor to market prices rather than optimistic midpoints to produce credible estimates SEC Rule 605 modernization.
Midpoint fills are attractive because they are simple to compute and remove the immediate visible cost of the bid-ask spread, but they assume either perfect price improvement or the absence of queue friction. That assumption is optimistic in many real trading contexts because tick-size constraints, queue priority and hidden liquidity mean that executable prices diverge from midpoints, especially for larger orders or during volatile windows.
Commonly used idealizations persist in academic papers and hobbyist models because of limited access to consolidated, timestamped venue quotes and the engineering effort required to model slippage. The gap between an idealized simulation and live execution shows up as overstated alpha and understated risk unless the backtest explicitly models spreads, slippage and fees, and reconciles simulated trades against authoritative quote feeds.
Why Market Prices Should Be Included in Backtests: regulatory drivers and implications
The 2024 modernization of Rule 605 expanded execution-quality disclosures that matter directly to backtests, including spread statistics, price improvement and order-size details by venue; using those outputs helps calibrate realistic, venue-aware cost models SEC Rule 605 modernization and see the SEC statement SEC statement.
Regulation NMS amendments in 2024 also changed minimum pricing increments and increased transparency around better-priced orders and order competition, making midpoint-based fills systematically optimistic in environments where tick size and queue position affect execution probability Regulation NMS amendments.
Record precise decision timestamps; snapshot authoritative quotes or NBBO; apply size- and volatility-conditioned slippage models; include fees and rebates; reconcile every simulated trade to quotes and venue outputs; and validate results with out-of-sample tests and multiple-testing adjustments.
Taken together, these regulatory changes raise the practical bar for backtests: prefer executable quotes and build probabilistic models of price improvement instead of assuming midpoint execution. That shift reduces look-ahead optimism by tying simulated execution to disclosed venue outcomes and observable spread behavior.
Why Market Prices Should Be Included in Backtests: data alignment and execution benchmarks
Aligning decision timestamps to authoritative feeds such as the NBBO reduces consolidation and look-ahead errors when reconstructing the state of the market at the decision moment; timestamp misalignment creates artificial opportunities in the simulated record that would not exist against a live consolidated feed FCA Wholesale Data Market Study.
Implementation shortfall is a practical validation metric because it aggregates spread, slippage, fees and timing relative to the decision price, giving a single, comparable measure of execution quality that links strategy decisions to realized P&L The Implementation Shortfall: Paper versus Reality.
Using implementation shortfall as a checkpoint requires consistent choice of the decision price and consistent timestamp alignment to the chosen authoritative feed. That means a backtest should record the decision timestamp, snapshot the top-of-book or consolidated best quotes for that instant, and compute slippage and fees relative to that baseline.
A practical framework for executable backtests
Start from a clear, reproducible workflow: log each decision with a precise timestamp and intended size, snapshot the authoritative quote feed for that timestamp, apply a size- and volatility-conditioned slippage model, subtract explicit fees and rebates, then record the executed price and compute implementation shortfall. This stepwise approach makes simulated fills auditable and comparable to later execution-quality disclosures SEC Rule 605 modernization. See Rule 605 details.
Required data inputs include timestamped top-of-book venue quotes or consolidated NBBO snapshots, trade prints for depth reconstruction where available, fee schedules and venue-level execution reports. When exchange-level execution-quality outputs are available, use them to calibrate price-improvement probabilities and fee effects.
For outputs, produce an adjusted P&L net of costs, a per-trade reconciliation table that lists decision price, executed price, quote snapshot and fees, and a validation checklist that flags clock sync issues, missing quotes and out-of-sample period status. Those artifacts turn an abstract backtest into a reproducible execution audit. See Funded Plays.
Modeling spreads, slippage and fees: building realistic cost models
Start cost modeling with quoted spreads observed in top-of-book snapshots rather than midpoints. Quoted spread and effective spread differ: the effective spread measures the actual execution disadvantage relative to the midpoint, but using the quoted spread or the venue best bid and offer for the decision instant gives a direct, conservative baseline for expected crossing costs Regulation NMS amendments.
Slippage should be modeled conditionally: small retail-size orders often clear at top-of-book with limited slippage, while larger tickets commonly walk the book and face price impact that increases with size and market volatility. Empirical slippage models can be parametric or non-parametric, but they must be calibrated to venue-level outcomes or to measured execution-quality reports when available.
Fees and rebates should be explicit line items in per-trade P&L. Rule 605-style disclosures and venue reports can help estimate typical price improvement and the frequency of executed orders inside the spread, which materially affects net cost once fees and rebates are considered.
Data quality, timestamps and consolidation: practical checks
Validate clock synchronization between decision logs and market feeds. Without synchronized clocks, a decision timestamp may be compared to the wrong quote snapshot, creating a look-ahead artifact; heartbeat checks and latency monitoring are practical ways to detect and measure clock drift in production environments FCA Wholesale Data Market Study.
Reconstructing NBBO from fragmented or delayed feeds carries consolidation lag risk; where possible, use authoritative consolidated feeds or exchange-provided timestamps for the decision instant to avoid reconstructing a false best bid or offer from stale messages SEC Rule 605 modernization.
When full, low-latency feeds are unaffordable, adopt pragmatic mitigations: calibrate slippage with a held-out sample of traded fills, run vendor parity checks to detect systematic biases between consolidated and vendor feeds, and document licensing limits so stakeholders understand data gaps.
Open execution-aware validation templates on the FundedPlays Challenges page
If you want a short validation checklist to apply these timestamp and consolidation checks, download or view the implementation templates designed for execution-aware backtests.
Document data lineage: store raw feed messages, transformation logs and reconciliation outputs. Keeping the raw feed makes later audits possible and enables replays when new venue disclosures or improved consolidation logic become available.
Validation and statistical safeguards: out-of-sample testing and overfitting controls
Realistic execution-aware backtests still need robust statistical safeguards. Use strict out-of-sample protocols such as time-based splits and walk-forward validation to test whether execution-aware gains persist, because in-sample performance can be inflated by parameter tuning and selection bias The Probability of Backtest Overfitting.
Apply multiple-testing corrections when evaluating many parameter combinations. Adjusted measures such as the deflated Sharpe help control false positives by accounting for the number of independent trials and the selection process that produced the best-performing configuration.
Always reconcile simulated fills against trade-level evidence where possible. A per-trade reconciliation that matches decision snapshots to executed prints provides an independent validation of the slippage assumptions and catches classes of errors that purely statistical tests might miss.
Common pitfalls and how to avoid them
Look-ahead bias is one of the most common mistakes. Examples include using consolidated prints that arrive after the decision time or using midpoints computed from future messages. Detect look-ahead by running negative-control tests that shuffle timestamps or by checking whether simulated trades consistently outperform venue-reported execution statistics The Probability of Backtest Overfitting.
Ignoring tick size and queue position also creates optimism. When minimum pricing increments and order priority affect whether a passive order executes, midpoint fills can be unavailable in practice; Regulation NMS changes emphasize that tick-size and queue effects are material to execution outcomes Order Competition Rule.
trade-level reconciliation checklist
keep raw feed records for audits
Underestimating venue effects and fee structures is another common trap. Test sensitivity to alternative fee schedules and to routing assumptions, and include fee lines in the reconciliation table so stakeholders can see how fees shift net returns. See how Funded Plays evaluations work.
Practical examples and scenarios
Scenario one, small retail-size orders: a prediction system that places small bets or small equity trade equivalents can often rely on top-of-book snapshots and a simple spread crossing cost. In this scenario, slippage is typically modest and can be estimated with a short calibration sample against venue reports, producing an adjusted P&L that mostly tracks the midpoint-based P&L but with a consistent per-trade cost deducted.
Scenario two, large institutional tickets: strategies that scale up to larger sizes need depth-of-book models. Large orders commonly walk the book or require passive posting with uncertain queue position, and execution quality depends on depth, hidden liquidity and queue priority; these effects can transform a strategy that looked profitable on midpoint fills into one with marginal net returns once depth and price impact are included The Implementation Shortfall: Paper versus Reality.
When presenting results, include per-trade reconciliation rows and aggregate metrics: average implementation shortfall, distribution of slippage by size bucket, and validation flags for data completeness. Those outputs help stakeholders understand when simulated performance is robust and when it depends on optimistic execution assumptions; see related blog posts related blog posts.
Checklist and next steps for teams adopting execution-aware backtests
Minimum acceptance checklist: verify clock sync across decision and feed logs; use NBBO-aligned or authoritative top-of-book quotes; calibrate a slippage model by size and volatility; implement per-trade reconciliation; and pass out-of-sample validation protocols SEC Rule 605 modernization.
Phased roadmap: start by applying top-of-book spread and explicit fee adjustments to existing backtests, then add venue-level slippage calibration using historical execution reports, and finally invest in depth-of-book and queue modeling as execution-quality disclosures and consolidated feeds become more accessible through 2025 and 2026 Regulation NMS amendments. See the implementation timeline implementation timeline.
Open gaps for further work include realistic modeling of hidden liquidity and queue position, which depend on instrument and venue microstructure beyond current disclosures. Prioritize incremental investments that improve the fidelity of reconciliation outputs and enable comparison to exchangedisclosed execution-quality statistics FCA Wholesale Data Market Study.
Midpoints and idealized fills ignore spreads, tick-size and queue effects, which typically understate costs. Using quoted prices or NBBO snapshots gives a more conservative, reproducible baseline.
Implementation shortfall measures the difference between the decision price and executed price, aggregating spread, slippage and fees. It is useful as a single, comparable execution-quality metric.
Use pragmatic workarounds: calibrate slippage on a sample of executed fills, run vendor parity checks and document data gaps. Incrementally improve models as better disclosures become available.
References
- https://www.federalregister.gov/documents/2024/05/14/2024-10776/order-execution-disclosures
- https://www.federalregister.gov/documents/2024/05/15/2024-10774/regulation-nms-minimum-pricing-increments-access-fees-and-transparency-of-better-priced-orders
- https://www.fca.org.uk/publications/market-studies/wholesale-data-market-study-final-report
- https://jpm.pm-research.com/content/14/3/21
- https://www.federalregister.gov/documents/2024/05/15/2024-10775/regulation-nms-order-competition
- https://papers.ssrn.com/sol3/papers.cfm?abstract_id=2326253
- https://www.fundedplays.com/challenges
- https://www.fundedplays.com
- https://www.fundedplays.com/blogs
- https://www.fundedplays.com/blogs/how-fundedplays-evaluations-work
- https://www.sec.gov/newsroom/speeches-statements/crenshaw-statement-order-execution-quality-030624
- https://clearstreet.io/legal/rule-605-606-detail
- https://flextrade.com/resources/sec-rule-605-is-final-but-more-is-pending-with-market-structure/
