From countless papers to strategies you can test.

Outsample reads the most relevant papers in your corpus and returns strategies, research programs, and answers backed by the papers it read. It stress tests its own ideas, and tells you when the evidence isn't there.

No card. 2 full synthesis runs, total. No expiry.

13,944 searchable papers in the corpus. 3 research programs for ~39 credits. 35 adversarial findings on the last strategy run.

01 / the bottleneck

The bottleneck is synthesis.

Reading fifty papers to find the three disagreements that matter takes weeks, and most ideas die in the backtest.

/ the same question, two answers

A frontier model answers with papers. Outsample answers with strategies you can test.

Outsample runs on Claude Sonnet. The other column is Claude Fable 5, a more capable model in the same family.

Asked of both, verbatim Could you provide 5 Volatility based strategies, based on citations off real research papers? Do not infer, use purely citation backed answers

Claude Fable 5, API, web search on, 2026-08-26
  1. Selling variance swaps / harvesting the variance risk premium Carr & Wu (2009), Variance Risk Premiums, RFS 22(3) Sells variance swaps on equity indexes, collecting the premium paid for volatility-spike insurance.
  2. Volatility-managed (inverse-variance-scaled) factor portfolios Moreira & Muir (2017), Volatility-Managed Portfolios, J. Finance 72(4) Scales factor exposure by the inverse of its recent realized variance.
  3. Low idiosyncratic volatility long/short equity Ang, Hodrick, Xing & Zhang (2006), The Cross-Section of Volatility and Expected Returns, J. Finance 61(1) Long the low idiosyncratic-volatility quintile, short the high.
  4. Trading the VIX futures basis (contango/backwardation roll) Simon & Campasano (2014), The VIX Futures Basis, J. Derivatives 21(3) Shorts VIX futures in contango, buys in backwardation, hedged with E-mini S&P futures.
  5. Volatility timing in portfolio allocation Fleming, Kirby & Ostdiek (2001), The Economic Value of Volatility Timing, J. Finance 56(1) Rebalances stock, bond, gold and cash weights daily from conditional covariance forecasts.
Its closing caveat

These are research findings. Nothing here claims the strategies still work. Its own close says as much: short-volatility positions carry large left-tail risk (the premium exists precisely as compensation for volatility-spike insurance), so realized performance outside the studied samples is not guaranteed by these papers.

Outsample answered from the 15 most relevant of 116 retrieved
  1. Variance Premium Decomposition Return Timing on SPY 3/5 SPY, VIX, SPX options Monthly: variance premium (VIX squared minus a GARCH forecast of realized variance) over 5 points goes 100% SPY, 50% when compressed.
  2. Volatility-Managed Momentum Factor via Realized Volatility Scaling 4/5 Russell 1000 long-short momentum, SPY overlay Weekly-rebalanced 12-1 decile momentum scaled to 1% daily vol, scalar capped 2x / floored 0.25x, flat exit after 3 weeks below the floor.
  3. HAR Realized Volatility Regime Timing on S&P 500 Index Options 3/5 VIX futures, SPX 30-day straddles HAR+options variance forecast vs implied; enter on a 2-point gap, 20-day hold, 1.5% NAV risk budget per trade.
  4. Cross-Market Volatility Spillover Timing: Equity-to-Oil Transfer Signal 2/5 WTI futures, USO, crude straddles SPX downside semi-variance spike above 400 points with an equity/oil ratio over 1.5 opens a crude straddle; 10-day hold, 60% profit target.
  5. Range-Based Volatility Breakout with GARCH Confirmation on FX Majors 3/5 EUR/USD, GBP/USD, USD/JPY 1-week straddles Yang-Zhang 5-day range vol over 1.25x GARCH(1,1) buys the straddle; hold to expiry, 80% profit target, 1.5% NAV per pair.

Every strategy gets a hostile review before you see it.

Its own adversarial review returned

abandon

That verdict is the product working. It graded these five 2 to 4 out of 5, then argued against them. A reading list has nothing to abandon.

/ added beyond the source papers

For every strategy it writes, Outsample lists which parameters came from the papers and which it invented. Run 43 disclosed 31 such extensions across the five. The competitor has nothing to disclose here, because it is not generating parameters. One strategy, in full:

Volatility-Managed Momentum Factor via Realized Volatility Scaling

  • the paperPaper tested volatility-managed portfolios on Fama-French factors and market portfolio. we addedApplying specifically to a self-constructed 12-1 momentum decile sort among Russell 1000 is our extension.
  • the paperThe paper tested target volatilities at different levels. we addedTarget volatility of 1% daily is our specification; we select a specific one here.
  • the paperThe paper used monthly rebalancing. we addedWeekly rebalance frequency is our extension.
  • the paperThe paper did not specify explicit bounds. we addedScalar cap of 2x and floor of 0.25x are our additions.
  • the paperNot present in the source paper. we addedFlat exit rule when scalar is below 0.25x for 3 consecutive weeks is our risk management extension.

Same request. One side returns papers about strategies; the other returns strategies you can test tomorrow, with every parameter it invented disclosed, and its own critic arguing against them. That last part is why one side can say abandon and the other has nothing to abandon.

02 / ask the corpus

Point it at a corpus. Ask. Get the answer and the papers it leaned on.

  • A synthesized answer, every claim cited to the paper it came from
  • The papers it leaned on, with their similarity scores
  • How many papers were retrieved, and how many it actually read
  • Flagged as thin when the corpus barely supports the question
  • The cost of the run, in frame
2026-08-04. Order-flow imbalance, answered from the corpus with every source cited. $0.06.
the workflow

How research ideas become vetted strategies

  1. Topic recommendations

    Give it a focus hint, or leave it blank to span your relevance scored corpus. Either way it produces research ideas worth looking into.

  2. Project proposals

    Take your pick of an idea, and it scaffolds a project around it, with a draft thesis, open questions, and seed papers to start from.

  3. Strategies

    From your project, it generates literature backed trade strategies based on the papers you selected, or the ones it pulled.

  4. Devil's Advocate

    Adversarially critiques its own strategies. Even its best rated strategies get stress tested before you pursue them.

Outsample Ask
‹ Corpus › Research › Ask
ask the corpus

Ask retrieves the most similar papers from the corpus and answers from the ones that clear the relevance gate, citing what it used. When too little clears, it says so.

Question 82/2000 Mode
Model Cost is the only difference: Budget is roughly a fifteenth of Standard, Max about twice. Paying more does not make it slower. Standard is the default.
The default every result to date has used.
50
≈ $0.07 · ~30s
Scope: AllSharedMy uploads
Sources: arxiveconstorrepecnberopenalexbismanual
Quick mode · 9 citations

What does the literature say about order flow imbalance as a short-horizon signal?

Answered from the 50 most relevant of 298 retrieved

Order flow imbalance (OFI) is one of the most studied short-horizon signals in the market microstructure literature, and the papers in this corpus collectively paint a rich picture of how it works, why it works, and what limits it.

Theoretical foundations: One paper provides an important theoretical grounding, showing that volume imbalance in limit order books is not merely a spurious empirical regularity but is an endogenously optimal response by informed market makers to price information. This work argues that the price-imbalance predictability observed empirically reflects rational behavior by market makers who adjust their quotes based on information about efficient prices, rather than a persistent market inefficiency that could be easily arbitraged away.

Empirical signal strength and horizon: The research consistently supports OFI as a genuine short-horizon return predictor across markets. One study on CSI 500 constituent stocks proposes a Generalized Stationarized Order Flow Imbalance (log-GOFI) metric and reports , highlighting that how OFI is computed matters greatly for predictive power. A separate study on BankNifty index futures finds that filtering transient or short-lived orders from the limit order book modestly sharpens OFI-return correlations, but that trade-based OFI, which focuses on parent order flow rather than aggregate quote activity, produces stronger directional signals than time-based or count-based filtration schemes.

Normalization and signal extraction: A paper applying a matched-filter perspective to Korean equity order flows emphasizes that the correct normalization of OFI is not one-size-fits-all. Market-cap normalization is optimal for detecting signals from capacity-constrained institutional investors, while volume normalization is better suited to execution-algorithm flow. The empirical results are striking, with , illustrating that poorly matched normalization suppresses otherwise strong signals.

Retail versus institutional order imbalance: One study examines retail order imbalance (ROI) specifically and reaches a cautionary conclusion. The predictive power documented in an earlier seminal study for the 2010 to 2015 period has deteriorated markedly in 2016 to 2021, and the conclusions depend heavily on the methodology used to identify retail trades. Long-short strategies built on ROI are no longer profitable in the more recent sample, underscoring that order flow signals can decay as markets adapt and as the composition of order flow changes.

Toxicity and adverse selection: A related dimension of OFI as a signal concerns flow toxicity. The VPIN (Volume-Synchronized Probability of Informed Trading) metric quantifies order flow toxicity based on volume imbalance and trade intensity, and the research on it shows that high imbalance in volume time predicts short-term volatility spikes driven by adverse selection. This frames OFI not only as a directional signal for price movement but also as a risk indicator for liquidity provision.

Execution use cases: Beyond pure alpha generation, several papers demonstrate OFI's value within execution algorithms. One agent-based simulation study shows that incorporating order book imbalance signals into VWAP and TWAP-style execution improves performance in volatile markets and is robust to spoofing manipulation, suggesting OFI is actionable within realistic execution constraints. A study on optimal execution under endogenous order flow modeled via Hawkes processes validates on 110 NASDAQ stocks that accounting for self-exciting order flow dynamics improves execution outcomes, which implicitly relies on OFI-like features capturing feedback effects.

Liquidity withdrawal and early warning: One study introduces a Liquidity Withdrawal Index forecasting framework using Nasdaq market-by-order data, finding that order flow features including cancellation ratios can predict liquidity withdrawal 250 milliseconds to 5 seconds ahead. This is a closely related application of OFI logic, where the imbalance between order arrivals and cancellations serves as an early-warning signal for order book stress rather than for directional price moves.

Overall picture: Research supports OFI as a robust short-horizon signal, but with important qualifications. Signal strength depends on how imbalance is constructed (raw versus filtered, what normalization is used, whether trades or quotes are the basis), on the trader type whose flow is being measured, on the market and instrument, and on the time period. A key theoretical insight is that the predictability is not a free lunch but reflects informed order routing, meaning it is competed away as more participants exploit it, as the retail order imbalance evidence illustrates. The most practically robust applications appear to combine careful signal construction with its use in execution optimization rather than as a standalone directional alpha source.

Cited papers

2026-08-26. Quick mode, $0.0758, 9 papers cited.

03 / the obvious question

My LLM already does this.

It doesn't, and the difference is mechanical. A chatbot answers from what it remembers or what a search returned, and it writes the citations with the same confidence as everything else.

Here, retrieval is its own step. Keyword and vector search run together, then a relevance gate cuts what does not bear on the question. On the disagreement run, 295 papers were recalled and 50 cleared the gate. What survives is what the synthesis reads, and every claim traces back to one of those papers.

Then it argues with itself before you see the result.

The test that settles it takes one run on a topic you know cold.

honest numbers

What it holds, and what it costs.

15,139 papers in the corpus, every one read for synthesis. 4,818 are shown in full: the papers whose licence is provably open to redistribute. The rest are read the same way and shown as reference entries.

Every paper is tagged at ingest with the themes it works in. The counts are papers per theme, measured from the shared library.

171 themes with thirty or more papers each
market microstructure 384, causal inference 433, portfolio optimization 267, behavioral finance 315
game theory 186, sentiment analysis 155, risk management 128, reinforcement learning 114.

Check your niche

A corpus question costs about 76 credits, a full strategy run around 420 credits. Every answer shows its exact cost.

Corpus figures measured 2026-08-24.

Pricing

Free

$0

25 papers, 5 corpus questions a month, 2 full synthesis runs, total. No card.

Pro

$39 / month

1,000 papers. About 36 runs a month on the included 15,000 credits at Standard rates, or about 197 lighter corpus questions.

Ultimate

$70 / month

10,000 papers, 45,000 credits included monthly. Deep runs get priority in the queue.

Team

$129 / month

Five seats, 10,000 papers, 60,000 credits included monthly. Set up per desk, by contact.

Annual billing is two months free. Model usage beyond included credit is prepaid and itemized, with a cost preview before every heavy run. Every query shows what it cost, and nothing bills past credit you bought. Full details: /pricing

From the builder

Before this, my research process was a folder of SSRN PDFs and a rough reading plan. Ten or twenty pages on a good day, and half of that was background reading just to understand a paper's premise. After weeks of that produced one strategy, and the strategy failed, I accepted that reading hundreds of papers one at a time was never going to find my ideas.

So I built the tool I wanted: I generate research programs off my corpus, expand the ones worth expanding, trace which papers connect, and find the central ones and the niche ones I would have missed.

It runs on a model and on the corpus you give it. I can't promise you a six-figure strategy. It makes the search faster and wider. The judgment is still yours.

Mario

Outsample. Because in-sample results lie.

Start free. No card. 2 full synthesis runs, total, and the product will tell you if your question is not answerable from the corpus.