Begin with a claim that can fail
A useful investment hypothesis describes a relationship, the observations that support it, and the evidence that would change the conclusion. A fluent explanation alone is insufficient. The data must have been available at the time the decision would have been made.
Our research framework separates generating an idea from evaluating it. A reviewer should be able to identify the source, sample period, transformations and assumptions without reconstructing a conversation.
Preserve a comparison
A backtest is an experiment with many opportunities for accidental optimism. Costs, allocation rules and a relevant benchmark belong in the experiment from the beginning. A period used to choose a model cannot also serve as independent evidence that the model works.
Keep a frozen baseline and reserve observations for evaluation. If a revision improves one metric while weakening robustness, turnover or drawdown, record that trade-off. A failed test is a useful result, not a reason to change the scorecard after seeing the outcome.
Promotion is a separate decision
Historical evaluation, real-time simulation and broker paper execution answer different questions. The first studies an idea; the second tests behavior as information arrives; the third rehearses order handling. None is equivalent to live investor performance.
SumTwo is developing this process around US stocks and ETFs. We are establishing operational evidence and a verifiable record. No research note on this site is a recommendation to transact or a claim that a strategy has demonstrated an investment edge.
