研究迴圈報告 — Run #13

研究主題:credit spreads · 產生時間:2026-03-11 10:01:43 UTC

總輪次
5
總假說數
64
已確認
0
已拒絕
0
待驗證
64
最高信心
0.9900

執行摘要

本次研究迴圈共進行 5 輪, 產生 64 個假說, 其中 0 個通過驗證、 0 個被拒絕、 64 個待驗證。

總耗時 0 秒, 消耗 0 tokens。

各輪摘要

輪次Agent 數Agents假說數耗時Tokens
R13optimizer, researcher, risk-auditor 15
R23optimizer, researcher, risk-auditor 16
R33optimizer, researcher, risk-auditor 12
R42devil, risk-auditor 10
R52portfolio-mgr, risk-auditor 11

假說清單

Round 1 (15 假說)

狀態維度信心假說
pending credit spreads 0.7500 Corrected methodology (FIRST-bar pricing, transaction costs, NaN-as-NaN) will show credit spreads are marginal or unprofitable vs Run 11's inflated results
pending credit spreads 0.6000 Iron condors (simultaneous put + call credit spreads) double premium collection but also double tail risk and transaction costs
pending credit spreads 0.5500 VIX-based entry filter (enter only on high-vol days) improves risk-adjusted returns by capturing higher premiums
pending credit spreads 0.9200 Bid-ask spreads for OTM puts are prohibitively wide, making realistic execution costs 10-30% of option premium for strikes under $5
pending credit spreads 0.9500 Only 3-5% of minute bars have tradeable prices (close>0), confirming severe data sparsity that makes minute-bar backtesting unreliable
pending credit spreads 0.9300 5% OTM put spread credits are near-zero or NEGATIVE 40-45% of the time using first-bar pricing, making the strategy unviable before costs
pending credit spreads 0.9000 Credit spread profitability is entirely concentrated in 2020-2021 (post-COVID high-vol); 2022-2025 shows zero or negative edge
pending credit spreads 0.7200 September and May are consistently the worst months for put credit spreads due to seasonal vol expansion events
pending credit spreads 0.8800 Average volume of 8-9 contracts per bar at 5% OTM makes real execution of institutional-size positions impossible
pending credit spreads 0.9400 Post-cost Sharpe ratio of 0.09 (annualized) confirms zero statistical edge in 5% OTM credit spreads
pending credit spreads 0.7200 VIX條件式鐵禿鷹策略(VIX>20時進場,30-45 DTE,雙側10-delta,$25寬翅膀,50%利潤目標管理)可實現Sharpe>0.8。Run 11確認VIX>20子樣本(2022年)的put側Sharpe=1.28、勝率80%。加入CALL側可收取雙倍權利金,擴大盈虧平衡區間。VIX條件過濾排除低波動期的微薄權利金問題(Run 11發現3% OTM在低波動時僅$0.05)。
pending credit spreads 0.5800 CALL信用價差(熊市CALL價差)在上漲趨勢市場中比PUT信用價差更安全,因為SPX的漲幅具有自然天花板(日漲幅極少超過3%),而跌幅可達-10%以上。以3% OTM短腿、$25寬的CALL價差在weekly(4-5 DTE)可獲得$1-4的credit,且被穿透的概率低於PUT側。
pending credit spreads 0.6500 動態止損+提前平倉規則可以將Run 11的OOS Sharpe從0.24提升至>0.6。具體規則:當價差虧損達到最大虧損的50%時平倉($25寬價差在浮虧$12.50時止損),配合50%利潤目標。此規則截斷左尾(Run 11的主要缺陷),代價是降低勝率但大幅改善風險調整後收益。
pending credit spreads 0.6000 30-45 DTE信用價差比weekly(4-7 DTE)有更好的風險調整後收益,因為:(1)更長的DTE提供更高的絕對premium(數據顯示60 DTE 3% OTM CALL=$50+ vs 4 DTE=$2),(2)theta曲線在20-30天時最陡峭,可在剩餘15 DTE時平倉捕獲60-70%的decay,(3)gamma風險在長DTE時較低,被穿透時的損失更漸進而非二元跳崖式。
pending credit spreads 0.5200 結合所有改進的「複合條件」策略——VIX>20進場 + 鐵禿鷹結構 + 30-45 DTE + 動態止損(50%虧損)+ 50%利潤目標——可實現OOS Sharpe>0.5且最大回撤<15%。此策略解決Run 11的所有已知缺陷:低premium(VIX過濾+longer DTE)、單邊風險(鐵禿鷹)、尾部虧損(止損)、hold-to-expiration decay(利潤目標)。

Round 2 (16 假說)

狀態維度信心假說
pending entry_timing 0.8500 Earlier entry windows (9:31-9:35) capture higher premiums than later windows, resulting in better avg P&L per trade
pending entry_timing 0.8000 The BASELINE window (9:31-9:45) produces the most trades (243 vs 165-198) and highest total P&L, making it the most robust entry approach
pending entry_timing 0.7000 5% OTM put spreads (5 wide) with corrected methodology show 96-98% win rate and Sharpe >5 across all entry windows, but these results may reflect 0DTE same-day expiration bias
pending entry_timing 0.9500 Transaction costs (.60/spread = 0.026 pts) are negligible relative to credits of 20-27 pts, contributing <0.1% drag
pending entry_timing 0.9200 Opening 15-min bar prices are significantly inflated vs intraday average, creating adverse selection for put sellers entering at open
pending entry_timing 0.8800 Intrabar slippage at market open is 60-80% larger than midday, eroding the apparent open-price advantage
pending entry_timing 0.9500 85% of OTM put contracts have NO first trade until after 9:30, with median first-trade delay of 28 minutes
pending entry_timing 0.9300 Cheap options (<$0.50) have 11.5% average intrabar spread—slippage alone can exceed the premium collected
pending entry_timing 0.9000 Volume peaks at open but remains sparse throughout the day—median volume is ZERO across all time windows for the full strike range
pending entry_timing 0.6500 2024 data shows convergence of open vs midday spreads, suggesting 0DTE growth has improved opening liquidity
pending entry_timing 0.8200 EOD prices show negligible bias vs midday (0.09%), but open prices are 64.5% higher—systematic intraday mean-reversion in OTM puts
pending entry_timing 0.9200 0DTE put credit spreads have 15x more volume and 5x more tradeable spread pairs than 1-5DTE, making 0DTE the only viable DTE for credit spread entry
pending entry_timing 0.7800 Entry at 10:30-13:00 provides best fill quality for 0DTE credit spreads — opening 30min has 15-53% wider high-low spreads vs midday
pending entry_timing 0.5500 Tuesday and Wednesday have highest credit-to-width ratios for 0DTE spreads; Friday has worst
pending entry_timing 0.8500 Volume>=10 filter on both legs is necessary but still leaves credit spreads with high execution risk due to 12-19% average high-low spread
pending entry_timing 0.8000 Overall assessment: entry timing optimization can improve credit spread fills by ~20% but cannot overcome the fundamental problem of thin edge + high execution costs

Round 3 (12 假說)

狀態維度信心假說
pending strike_selection 0.9900 R2 backtest results are INVALID due to a critical date-filter bug: the script queries parquet files by TIME only (e.g., 09:31-09:35) without filtering by DATE. Each SPXW_YYYYMMDD.parquet file contains ~40 days of historical price data for options expiring on that date. FIRST(close ORDER BY ts) picks...
pending strike_selection 0.9900 The $26.42 average credit, 98% win rate, Sharpe 6.12, and -$2.40 max loss are ALL artifacts of the date-filter bug and contain zero valid signal. Credit > spread width ($79 credit on $25 spread) is physically impossible for a credit spread — this alone proves the data is corrupted. Every metric in t...
pending strike_selection 0.7000 The spot estimation via put-call parity is also contaminated by the same date-filter bug but happens to produce a correct result by coincidence. The ATM strike selection query also lacks a date filter, pulling prices from prior days, but the closest put-call parity match at ATM strikes tends to be f...
pending strike_selection 0.9500 Secondary issue: Sharpe ratio uses sqrt(52) weekly annualization, but 0DTE strategies trade daily (~250 days/year). This would overstate Sharpe by sqrt(252)/sqrt(52) = 2.2x even if P&L data were correct.
pending strike_selection 0.9200 R2 avg credit of $26.42 on 5% OTM 0DTE $25-wide put spread is an artifact of the spot estimation bug — put-call parity on 0DTE returns the ATM STRIKE (not spot), making the 'spot' estimate equal to the strike where |call-put| is minimized, which IS near ATM. But the 5% OTM target then selects strike...
pending strike_selection 0.9500 5% OTM 0DTE SPX put premium should be $0.05-$0.50 (0.5-5 cents in SPX points), not $26 — academic and practitioner data confirm deep OTM 0DTE options are nearly worthless
pending strike_selection 0.7800 Put-call parity spot estimation is unreliable for 0DTE options due to extreme bid-ask spreads, early exercise premium absence (European-style SPX), and dividend/rate effects becoming noise-dominant at sub-1-day horizons
pending strike_selection 0.8200 Delta-based strike selection (e.g., sell 10-delta put) is superior to percentage-OTM for credit spreads because it automatically adjusts for IV regime — in high-IV environments, 10-delta is farther OTM; in low-IV, it's closer
pending strike_selection 0.7500 Spread width of $25 is suboptimal — $5 width has better Sharpe per unit risk, while $50 width has better absolute returns but catastrophic tail risk on 0DTE
pending strike_selection 0.9300 The R2 results showing max loss of only -$2.40 on a $25-wide spread confirms a critical bug: either exit prices are wrong (NaN treated as 0 = expired worthless even when ITM) or the strikes selected are so far OTM that they never go ITM
pending strike_selection 0.7000 ATM-relative (moneyness) strike selection with fixed-dollar spread width is the most practical approach for backtesting, but must use the nearest standard 5-point strike grid rather than arbitrary interpolation
pending strike_selection 0.8800 The entire 0DTE 5% OTM put credit spread strategy is likely unprofitable after realistic execution costs — the premium is too thin ($0.05-$0.50) relative to bid-ask spreads ($0.05-$0.20) and the $2.60 round-trip commission

Round 4 (10 假說)

狀態維度信心假說
pending expiration_choice 0.8500 0DTE iron condor profitability varies significantly by OTM distance and width; wider/further OTM spreads may sacrifice win rate for better risk-adjusted returns
pending expiration_choice 0.8000 Closer OTM iron condors have higher win rates but worse risk/reward ratios due to smaller credit relative to max loss
pending expiration_choice 0.9000 Tail risk in 0DTE iron condors is substantial — max drawdowns can exceed total cumulative profit, requiring strict position sizing
pending expiration_choice 0.1500 1% OTM 0DTE put credit spread has avg credit of $1.26 (prior R3 finding)
pending expiration_choice 0.0200 1% OTM 0DTE spread is viable with post-cost Sharpe 0.09
pending expiration_choice 0.9800 Prior backtests used close price for exit valuation
pending expiration_choice 0.0500 Stress test: 1% OTM spread survives big down days
pending expiration_choice 0.2000 0DTE is the optimal DTE for credit spreads
pending expiration_choice 0.1000 Survivorship bias: missing expiration dates inflate results
pending expiration_choice 0.0500 SPX 0DTE credit spreads are viable in ANY configuration

Round 5 (11 假說)

狀態維度信心假說
pending position_sizing 0.9200 Kelly criterion sizing for 2% OTM / $25 wide config yields sub-optimal risk-adjusted returns versus passive alternatives
pending position_sizing 0.8800 Estimated annual return on $100K is 1.4% gross, negative after costs — economically unviable
pending position_sizing 0.9500 2022-2023 bear market data proves strategy has catastrophic left-tail risk that 2024 bull market conceals
pending position_sizing 0.9700 Data quality is fundamentally insufficient to validate ANY options selling strategy
pending position_sizing 0.9700 FINAL VERDICT: NO-GO for SPX credit spread deployment. Confidence: 97%. Score: 1.5/10
pending position_sizing 0.9700 ITM close=0 bug understates real losses by 70-95% on losing trades, making ALL reported win rates and P&L figures unreliable upper bounds
pending position_sizing 0.9500 Real-world P&L after bid-ask slippage, ITM bug correction, and commissions renders ALL tested configurations net-negative in expectation
pending position_sizing 0.9300 Risk of ruin exceeds 50% at any position size above 2% of capital for all tested configurations, due to max-loss-to-avg-win asymmetry
pending position_sizing 0.9200 Final risk rating: 2.0/10 — DETERIORATED from Run 11's 2.5/10 due to newly quantified ITM close=0 impact and iron condor call-side adding complexity without improving risk-adjusted returns
pending position_sizing 0.9000 Comparison to Run 11: situation has WORSENED — Run 13 added iron condor call side but the additional complexity produced no improvement, while deeper data analysis revealed the ITM close=0 bug is more severe than previously understood
pending position_sizing 0.9600 CONVERGENCE VERDICT: SPX 0DTE credit spreads on this dataset are NOT viable for live trading. The research program has exhausted what can be learned from this data. Further iteration without new data sources will produce diminishing returns.

錯誤日誌摘要

最近錯誤

方法評分

Agent方法成功率提出數確認數樣本數
optimizercredit spreads0.00%303
risk-auditorcredit spreads0.00%707
researchercredit spreads0.00%505
optimizerentry_timing0.00%404
risk-auditorentry_timing0.00%707
researcherentry_timing0.00%505
risk-auditorstrike_selection0.00%404
researcherstrike_selection0.00%808
risk-auditorexpiration_choice0.00%303
devilexpiration_choice0.00%707
portfolio-mgrposition_sizing0.00%505
risk-auditorposition_sizing0.00%606

建議與後續行動