Outsample reads the most relevant papers in your corpus and returns
strategies, research programs, and answers backed by the papers it
read. It stress tests its own ideas, and tells you when the evidence
isn't there.
13,944 searchable papers in the corpus.3 research programs for ~39 credits.35 adversarial findings on the last strategy run.
01 / the bottleneck
The bottleneck is synthesis.
Reading fifty papers to find the three disagreements that matter takes weeks, and most ideas die in the backtest.
▮ / the same question, two answers
A frontier model answers with papers. Outsample answers with strategies you can test.
Outsample runs on Claude Sonnet. The other column is Claude Fable 5, a more capable model in the same family.
Asked of both, verbatimCould you provide 5 Volatility based strategies, based on citations off real research papers? Do not infer, use purely citation backed answers
ClaudeFable 5, API, web search on, 2026-08-26
Selling variance swaps / harvesting the variance risk premiumCarr & Wu (2009), Variance Risk Premiums, RFS 22(3)Sells variance swaps on equity indexes, collecting the premium paid for volatility-spike insurance.
Volatility-managed (inverse-variance-scaled) factor portfoliosMoreira & Muir (2017), Volatility-Managed Portfolios, J. Finance 72(4)Scales factor exposure by the inverse of its recent realized variance.
Low idiosyncratic volatility long/short equityAng, Hodrick, Xing & Zhang (2006), The Cross-Section of Volatility and Expected Returns, J. Finance 61(1)Long the low idiosyncratic-volatility quintile, short the high.
Trading the VIX futures basis (contango/backwardation roll)Simon & Campasano (2014), The VIX Futures Basis, J. Derivatives 21(3)Shorts VIX futures in contango, buys in backwardation, hedged with E-mini S&P futures.
Volatility timing in portfolio allocationFleming, Kirby & Ostdiek (2001), The Economic Value of Volatility Timing, J. Finance 56(1)Rebalances stock, bond, gold and cash weights daily from conditional covariance forecasts.
Its closing caveat
These are research findings. Nothing here claims the strategies still work. Its own close says as much: short-volatility positions carry large left-tail risk (the premium exists precisely as compensation for volatility-spike insurance), so realized performance outside the studied samples is not guaranteed by these papers.
▮ Outsampleanswered from the 15 most relevant of 116 retrieved
Variance Premium Decomposition Return Timing on SPY3/5SPY, VIX, SPX optionsMonthly: variance premium (VIX squared minus a GARCH forecast of realized variance) over 5 points goes 100% SPY, 50% when compressed.
Volatility-Managed Momentum Factor via Realized Volatility Scaling4/5Russell 1000 long-short momentum, SPY overlayWeekly-rebalanced 12-1 decile momentum scaled to 1% daily vol, scalar capped 2x / floored 0.25x, flat exit after 3 weeks below the floor.
HAR Realized Volatility Regime Timing on S&P 500 Index Options3/5VIX futures, SPX 30-day straddlesHAR+options variance forecast vs implied; enter on a 2-point gap, 20-day hold, 1.5% NAV risk budget per trade.
Cross-Market Volatility Spillover Timing: Equity-to-Oil Transfer Signal2/5WTI futures, USO, crude straddlesSPX downside semi-variance spike above 400 points with an equity/oil ratio over 1.5 opens a crude straddle; 10-day hold, 60% profit target.
Range-Based Volatility Breakout with GARCH Confirmation on FX Majors3/5EUR/USD, GBP/USD, USD/JPY 1-week straddlesYang-Zhang 5-day range vol over 1.25x GARCH(1,1) buys the straddle; hold to expiry, 80% profit target, 1.5% NAV per pair.
Every strategy gets a hostile review before you see it.
Its own adversarial review returned
abandon
That verdict is the product working. It graded these five 2 to 4 out of 5, then argued against them. A reading list has nothing to abandon.
→ / added beyond the source papers
For every strategy it writes, Outsample lists which parameters came from the papers and which it invented. Run 43 disclosed 31 such extensions across the five. The competitor has nothing to disclose here, because it is not generating parameters. One strategy, in full:
Volatility-Managed Momentum Factor via Realized Volatility Scaling
the paperPaper tested volatility-managed portfolios on Fama-French factors and market portfolio.we addedApplying specifically to a self-constructed 12-1 momentum decile sort among Russell 1000 is our extension.
the paperThe paper tested target volatilities at different levels.we addedTarget volatility of 1% daily is our specification; we select a specific one here.
the paperThe paper used monthly rebalancing.we addedWeekly rebalance frequency is our extension.
the paperThe paper did not specify explicit bounds.we addedScalar cap of 2x and floor of 0.25x are our additions.
the paperNot present in the source paper.we addedFlat exit rule when scalar is below 0.25x for 3 consecutive weeks is our risk management extension.
Same request. One side returns papers about strategies; the other returns strategies you can test tomorrow, with every parameter it invented disclosed, and its own critic arguing against them. That last part is why one side can say abandon and the other has nothing to abandon.
02 / ask the corpus
Point it at a corpus. Ask. Get the answer and the papers it leaned on.
A synthesized answer, every claim cited to the paper it came from
The papers it leaned on, with their similarity scores
How many papers were retrieved, and how many it actually read
Flagged as thin when the corpus barely supports the question
The cost of the run, in frame
2026-08-04. Order-flow imbalance, answered from the corpus with every source cited. $0.06.
the workflow
How research ideas become vetted strategies
Topic recommendations
Give it a focus hint, or leave it blank to span your relevance scored corpus. Either way it produces research ideas worth looking into.
Project proposals
Take your pick of an idea, and it scaffolds a project around it, with a draft thesis, open questions, and seed papers to start from.
Strategies
From your project, it generates literature backed trade strategies based on the papers you selected, or the ones it pulled.
Devil's Advocate
Adversarially critiques its own strategies. Even its best rated strategies get stress tested before you pursue them.
▮ Outsample·Ask
‹ Corpus › Research › Ask
ask the corpusHistory
Ask retrieves the most similar papers from the corpus and answers from the ones that clear the relevance gate, citing what it used. When too little clears, it says so.
Question82/2000Mode
ModelCost is the only difference: Budget is roughly a fifteenth of Standard, Max about twice. Paying more does not make it slower. Standard is the default.
The default every result to date has used.
50
5how many papers the answer is written from100
✓ Force refresh (skip cache)≈ $0.07 · ~30sAsk
Scope:AllSharedMy uploads
Sources:arxiveconstorrepecnberopenalexbismanual
Quick mode · 9 citations☆ SaveShare
What does the literature say about order flow imbalance as a short-horizon signal?
Answered from the 50 most relevant of 298 retrieved
Order flow imbalance (OFI) is one of the most studied short-horizon signals in the market microstructure literature, and the papers in this corpus collectively paint a rich picture of how it works, why it works, and what limits it.
Theoretical foundations: One paper provides an important theoretical grounding, showing that volume imbalance in limit order books is not merely a spurious empirical regularity but is an endogenously optimal response by informed market makers to price information. This work argues that the price-imbalance predictability observed empirically reflects rational behavior by market makers who adjust their quotes based on information about efficient prices, rather than a persistent market inefficiency that could be easily arbitraged away.
Empirical signal strength and horizon: The research consistently supports OFI as a genuine short-horizon return predictor across markets. One study on CSI 500 constituent stocks proposes a Generalized Stationarized Order Flow Imbalance (log-GOFI) metric and reports , highlighting that how OFI is computed matters greatly for predictive power. A separate study on BankNifty index futures finds that filtering transient or short-lived orders from the limit order book modestly sharpens OFI-return correlations, but that trade-based OFI, which focuses on parent order flow rather than aggregate quote activity, produces stronger directional signals than time-based or count-based filtration schemes.
Normalization and signal extraction: A paper applying a matched-filter perspective to Korean equity order flows emphasizes that the correct normalization of OFI is not one-size-fits-all. Market-cap normalization is optimal for detecting signals from capacity-constrained institutional investors, while volume normalization is better suited to execution-algorithm flow. The empirical results are striking, with , illustrating that poorly matched normalization suppresses otherwise strong signals.
Retail versus institutional order imbalance: One study examines retail order imbalance (ROI) specifically and reaches a cautionary conclusion. The predictive power documented in an earlier seminal study for the 2010 to 2015 period has deteriorated markedly in 2016 to 2021, and the conclusions depend heavily on the methodology used to identify retail trades. Long-short strategies built on ROI are no longer profitable in the more recent sample, underscoring that order flow signals can decay as markets adapt and as the composition of order flow changes.
Toxicity and adverse selection: A related dimension of OFI as a signal concerns flow toxicity. The VPIN (Volume-Synchronized Probability of Informed Trading) metric quantifies order flow toxicity based on volume imbalance and trade intensity, and the research on it shows that high imbalance in volume time predicts short-term volatility spikes driven by adverse selection. This frames OFI not only as a directional signal for price movement but also as a risk indicator for liquidity provision.
Execution use cases: Beyond pure alpha generation, several papers demonstrate OFI's value within execution algorithms. One agent-based simulation study shows that incorporating order book imbalance signals into VWAP and TWAP-style execution improves performance in volatile markets and is robust to spoofing manipulation, suggesting OFI is actionable within realistic execution constraints. A study on optimal execution under endogenous order flow modeled via Hawkes processes validates on 110 NASDAQ stocks that accounting for self-exciting order flow dynamics improves execution outcomes, which implicitly relies on OFI-like features capturing feedback effects.
Liquidity withdrawal and early warning: One study introduces a Liquidity Withdrawal Index forecasting framework using Nasdaq market-by-order data, finding that order flow features including cancellation ratios can predict liquidity withdrawal 250 milliseconds to 5 seconds ahead. This is a closely related application of OFI logic, where the imbalance between order arrivals and cancellations serves as an early-warning signal for order book stress rather than for directional price moves.
Overall picture: Research supports OFI as a robust short-horizon signal, but with important qualifications. Signal strength depends on how imbalance is constructed (raw versus filtered, what normalization is used, whether trades or quotes are the basis), on the trader type whose flow is being measured, on the market and instrument, and on the time period. A key theoretical insight is that the predictability is not a free lunch but reflects informed order routing, meaning it is competed away as more participants exploit it, as the retail order imbalance evidence illustrates. The most practically robust applications appear to combine careful signal construction with its use in execution optimization rather than as a standalone directional alpha source.
Cited papers
2026-08-26. Quick mode, $0.0758, 9 papers cited.
▮ Outsample·Disagreement
‹ Corpus › Research › Disagreement
disagreementHistory
The model identifies paper-vs-paper disagreements in a topic area: opposing positions, the axis they differ on, and the likely explanation for the split.
Topic93/2000Mode
Model
Strongest reasoning; ~2x Standard's cost.
Papers retrieved50
5how many papers the answer is written from100
Force refresh (skip cache)≈ $0.13 · ~30sFind disagreements
Quick mode · 7 citations · 3 pairs☆ SaveShare
Does cross-sectional momentum decay or persist? Where does the literature genuinely disagree?
Answered from the 41 most relevant of 293 retrieved
Send to Combiner
Likely explanation
Axisprimary driver of stock momentum: industry grouping versus factor-return autocorrelation
Paper ADo Industries Explain Momentum?Claims industry momentum explains a substantial portion of the individual-stock momentum anomaly and that industry-level strategies are more profitable than stock-level strategies.paper_id: 160093
Paper BFactor Momentum and the Momentum FactorClaims individual stock momentum is fundamentally driven by momentum in factor returns and their autocorrelation structure, locating the source at the factor level rather than the industry level.paper_id: 160102
Likely explanation
Different unit of aggregation and methodology: one attributes momentum to industry classification, the other to latent factor momentum, so they compete over which structure subsumes stock-level momentum.
Axiswhether momentum is an industry-level cross-sectional effect or an individual-stock time-series effect
Paper ADo Industries Explain Momentum?Locates the origin of momentum in industry-level return continuation, implying stock momentum is largely an industry effect.paper_id: 160093
Paper BCross-Sectional and Time-Series Determinants of Momentum ReturnsLocates the origin in time-series autocorrelation of individual stock returns rather than in any cross-sectional grouping such as industries.paper_id: 160327
Likely explanation
Methodology choice over decomposition: industry-sorted portfolios versus a cross-sectional/time-series variance decomposition of individual stock returns lead to different attributions.
Cited papers
2026-08-26. Quick mode on Max (Opus 4.8), $0.1353, 7 citations, 3 disagreement pairs.
▮ Outsample·Topic recommendations
‹ Corpus › Research › Topic recommendations
topic recommendationsHistory
The model picks focus topics worth investigating, grounded in your project's literature (or a corpus hint). Each card routes into Ideas / Strategies or Combiner with one click.
Focus hint88/500Model
The default every result to date has used.
How many topics7
312
≈ $0.14–$0.21 · ~60sRecommend topics
7 topics · 100 papers used
Recommended from 100 papers of 576 retrieved
Hint: Intraday momentum and microstructure-based short-horizon signals on equity index futures
#1
Papers:
Use in:
#2
Papers:
Use in:
#3
Papers:
Use in:
2026-08-26. $0.1256, 7 topics from 100 papers.
▮ Outsample·Strategies
‹ Corpus › Research › Strategies
strategiesHistory
The model retrieves the most-similar papers in your corpus and proposes concrete trading strategies you can test or combine.
This strategy captures the well-documented momentum premium while dynamically scaling position size using real-time forecasts of momentum's conditional mean and variance. The core insight from Barroso and Santa-Clara (paper 160189) and Daniel and Moskowitz (paper 160100) is that momentum risk is highly time-varying and predictable, concentrated in high-volatility panic states following market declines. By reducing gross exposure when predicted momentum variance is elevated and increasing it when variance is low, the strategy aims to eliminate the crash component while preserving the positive-return episodes. The scaling rule targets a constant ex-ante volatility for the momentum factor, normalizing by its realized variance over the prior 126-day (6-month) window.
Instruments
Signal
Entry
Exit
Position sizing
Expected behavior
Data required
Supporting papers
Export spec
Devil's Advocate7 flags · 4 high
Key test
Lookahead
The variance window and the VIX check are ambiguous about whether they close before or at the rebalance close; as written, the final day's values may not be knowable at decision time.
Overfit
Regime
The protection assumes variance rises before crashes. Post-2010 crowded momentum unwinds can spike variance simultaneously with the crash, and a 126-day window will lag a sudden regime break.
Costs
Scale
Manageable at $1M to $50M. Above roughly $200M the short leg's monthly rebalance becomes capacity-constrained and impact costs grow nonlinearly.
Implementation
Replicating UMD live needs CRSP-equivalent data, listing status, and the NYSE median-cap breakpoint, none of which come from standard broker feeds.
Statistical validity
2026-08-26. $0.4172, 5 strategies, 35 adversarial findings. The most expensive tool.
▮ Outsample·Project proposals
Projects › Project proposals
suggest a projectHistory
The model proposes full project shapes (draft thesis, seed papers, research questions) from corpus content. Click "Create project" on a proposal to materialize it as a real workspace with the seed papers queued for relevance scoring.
Focus hint (optional)68/500
How many proposals3
15
15 papers
How many papers to feed the model after retrieval + filters. Higher = more breadth, more cost.
≈ $0.06–$0.10 · ~10–25sGenerate proposals
3 proposals · 15 papers considered
Hint: Short-horizon microstructure signals worth building a project around
#1
#2
#3
2026-08-26. $0.0394, 3 project proposals from 15 papers. The on-ramp: how a research project begins.
03 / the obvious question
My LLM already does this.
It doesn't, and the difference is mechanical. A chatbot answers from what it remembers or what a search returned, and it writes the citations with the same confidence as everything else.
Here, retrieval is its own step. Keyword and vector search run together, then a relevance gate cuts what does not bear on the question. On the disagreement run, 295 papers were recalled and 50 cleared the gate. What survives is what the synthesis reads, and every claim traces back to one of those papers.
Then it argues with itself before you see the result.
The test that settles it takes one run on a topic you know cold.
honest numbers
What it holds, and what it costs.
15,139 papers in the corpus,
every one read for synthesis.
4,818 are shown in
full: the papers whose licence is provably open to redistribute. The rest
are read the same way and shown as reference entries.
Every paper is tagged at ingest with the themes it works in. The
counts are papers per theme, measured from the shared library.
171 themes with thirty or more papers each
market microstructure 384, causal inference 433, portfolio optimization 267, behavioral finance 315
game theory 186, sentiment analysis 155, risk management 128, reinforcement learning 114.
A corpus question costs about 76 credits, a full strategy run
around 420 credits. Every answer shows its exact cost.
Corpus figures measured 2026-08-24.
Pricing
Free
$0
25 papers, 5 corpus questions a month, 2 full synthesis runs, total. No card.
Pro
$39 / month
1,000 papers. About 36 runs a month on the included 15,000 credits at Standard rates, or about 197 lighter corpus questions.
Ultimate
$70 / month
10,000 papers, 45,000 credits included monthly. Deep runs get priority in the queue.
Team
$129 / month
Five seats, 10,000 papers, 60,000 credits included monthly. Set up per desk, by contact.
Annual billing is two months free. Model usage beyond included credit is
prepaid and itemized, with a cost preview before every heavy run. Every
query shows what it cost, and nothing bills past credit you bought.
Full details: /pricing
▮ Outsample
From the builder
Before this, my research process was a folder of SSRN PDFs and a rough reading plan. Ten or twenty pages on a good day, and half of that was background reading just to understand a paper's premise. After weeks of that produced one strategy, and the strategy failed, I accepted that reading hundreds of papers one at a time was never going to find my ideas.
So I built the tool I wanted: I generate research programs off my corpus, expand the ones worth expanding, trace which papers connect, and find the central ones and the niche ones I would have missed.
It runs on a model and on the corpus you give it. I can't promise you a six-figure strategy. It makes the search faster and wider. The judgment is still yours.
Mario
Outsample. Because in-sample results lie.
Start free. No card. 2 full synthesis runs, total, and the product will tell you if your question is not answerable from the corpus.