Drawdown Focus: BTCUSDT:/USDT 1h Analysis | Freqtrade | Kiploks

WHY THIS MATTERS

Backtests show what worked.

Kiploks shows what can survive.

This page answers one question:

"Can I safely deploy capital here?"

Strategy:OKXSmartGridFreqAI (Freqtrade)26 Jun 2026, 07:35
Asset:BTCUSDT:/USDT·Timeframe:1h·Exchange:okx
Test Period:2020-01-012025-04-25
Freqtrade config fee: 0.0015, symbol: BTCUSDT:USDT, endDate: 2025-04-25, exchange: okx, startDate: 2020-01-01, timeframe: 1h, initialBalance: 500, Trades: 1740

FINAL VERDICT

DO NOT DEPLOY

Diagnostic case: Neutral / Incubate (© Kiploks)

Kiploks Robustness Score: 0 / 100Bayesian pass probability: 72%

One or more hard gates failed. DO NOT DEPLOY until blocking modules are fixed.

Deployment Gate
Validation Gates
  • Execution Buffer - Net Edge (Net Profit > 15 bps, period-level)(-40.07 bps vs 15 bps) ESTIMATED - Net edge below 15 bps or edge deficit after fees
  • Stability (WFE > 0.5)(0.30 vs 0.5) - WFE below 0.5 (OOS/IS ratio too low)
  • Data Quality Guard (test period ≥ 2 years)(1941 vs 730 days)
Statistical Confidence
  • Statistical Significance (t-Stat > 1.96)
  • t-Stat (OOS Edge) > 2.0 (same metric as above, stricter threshold)(3.78 vs 2)
Critical Failures
  • Deployment is blocked because the following hard gate(s) failed: Execution Buffer - Net Edge (Net Profit > 15 bps, period-level); Stability (WFE > 0.5).
Execution Note: Backtest execution settings were missing. The system applied a standard Safety Buffer (0.05% slippage, 0.1% fee). For Institutional Grade (AAA), provide exact exchange API fees and liquidity-based slippage. You can set slippage and commission in your backtest or integration config to use exact values and remove this note.
Operational Insight: Net edge is negative; no slippage headroom. Status: Edge Deficit.

Robustness score is 0 because a module blocks (e.g. Risk, Execution, or Stability). Potential score if unblocked: 11. Fix blocking modules first. Even unblocked, score remains in TRASH range (0-20) - no meaningful improvement.

ROBUSTNESS SCORE

Data Quality: Outlier Influence 94%.
Formula (multiplicative penalties)
0 / 100 (FAIL)
Blocked by Walk-Forward & OOS, Execution Realism modules

Diagnosis

Execution score (5/100) is below the blocking threshold of 10. Edge does not survive 10 bps slippage - strategy may not be realizable in live conditions. Review transaction costs, reduce turnover, or improve edge.

░░░░░░░░░░░░░░░░░░░░
Methodology Note[?]
Breakdown (contributing factors)
Data Quality Guard
███████████░
94
→ Insufficient data period - overall score forced to Fail
Walk-Forward & OOS(40%)(blocking)
░░░░░░░░░░░░
0
→ BLOCKED
Risk Profile(30%)
████████░░░░
70
→ Risk metrics within tolerance
Parameter Stability(20%)
████████████
100
→ Parameters stable across sensitivity tests
Execution Realism(10%)(blocking)
░░░░░░░░░░░░
0
→ BLOCKED (raw 5, threshold 10)

DATA QUALITY GUARD

Data Quality Guard (DQG)
Score: 94%PASSDQG Factor: 0.94Contribution: 37.8

Robust Net Edge (Safe Edge): Profit is well distributed

Data Quality: Outlier Influence 94%.

ModuleScoreVerdict
Gap Density100%PASS
Outlier Influence94%PASS
Look-Ahead Bias100%PASS
Spread/Liquidity100%PASS
Sampling & Over-fitting100%PASS
Price Integrity100%PASS

BENCHMARK METRICS

Walk-Forward validation: [A] OOS-only metrics, [B] period-level (WFE, retention, profitable windows), [C] full backtest. Kill Switch and verdict use these values; at small N interpret with caution.
Quick Win (WFA summary)
[A] OOS equity-based OOS equity-based metrics from validation windows (canonical).From: WFA window OOS
OOS Sharpe: 0.68(same as below)OOS Calmar: 0.49OOS Max DD (validation only): -14.34%
[B] WFA period-levelWFA period-level: WFE, retention, profitable windows, trend match, win rate degradation.
WFE (Median OOS/IS): 0.30 (N=11 of 16, IS>0 only)OOS Retention: 46.0%
[C] Full backtest contextFull backtest (IS) context; canonical full_backtest.From: full backtest (no WFA split)
Full Sharpe: 0.08Full Calmar: 2.70
Profitable Windows[?]12/16 all windows (75%) PASS
Profitable OOS among IS>0 windows(same base as WFE)9/11
OOS/IS Trend MatchYES
Win Rate Change (OOS − IS, pp)[?]+1.4 pp (improvement)
IS Win Rate / OOS Win Rate62.9% / 64.3%
IS: full backtest; OOS: OOS trades
Statistical Robustness (OOS Validation)

WFE is conditional on IS>0 only - not a full-sample metric. Do not interpret WFE median as a full-strategy summary; the strategy may be loss-making overall. Min/median/max and variance use the same N (windows with IS > 0) (N=11): median is the middle value (odd N) or the average of the two middle values (even N). When overall OOS is negative, median WFE > 1 means the loss is driven by a subset of windows (in some windows OOS was better than IS). Large spread (min to max) with negative total OOS suggests a few windows or outliers dominate; interpret with caution. OOS Retention and Relative Change use all windows (N=16). Profitable counts use all windows. Retention over all windows reflects full P&L.

WFE (Median OOS/IS) distribution
min-0.40|median0.30|max1.16scale

Based on 11 windows with IS>0 (of 16 total). 2 windows had OOS<0.

One of the windows with IS>0 had OOS < 0 (see WFE min above). Median can be misleading with few windows and one strong negative; min shows risk of collapse.

WFE Variance0.248(population)(Consistency: Medium)
Parameter Stability Index (PSI)[?]n/a[?]
Edge Half-Life (T1/2, OOS)[?]2 windows (~240.0d)
WFA Windows16
Exp. OOS Return ± Vol[?]7.1% ± 10.4%
Worst Window Return (N=16, 1 obs)[?]
-14.34%
Avg OOS Sharpe (window-level)[?]0.68
Avg OOS Calmar (approx)[?]0.49
Optimization Gain (IS)[?]130.86%
OOS/IS Return Ratio[?]46.0%(N=16 windows)

sum(OOS)/sum(IS) over all windows(N=16 windows)

Relative Change (OOS−IS)/|IS|(mean OOS vs mean IS)[?]-54.0%

Negative = OOS worse than IS (degradation).

(mean(OOS)-mean(IS))/|mean(IS)|(N=16 windows)

Advanced Diagnostic Indicators
Verdict[?]
🔴REJECT

Immediate Kill Switch triggered. Net Edge < 10 bps (current: -40.07 bps); Consecutive OOS drawdown windows: 2 (limit: 1)

Capital Kill Switch - The Red Line[?]
2 consecutive (all windows) (limit: 1)

Next OOS window in minus → turn off bot

Summary (possible)
Regime Failure: Strategy passed only 1 of 3 market regimes (1/3 pass). Logic is not adapted to current market.
OOS/IS return ratio 46.0% indicates strong performance drop to OOS. Optimization gain (IS - OOS): 130.9%.
Statistical Confidence: Bayesian pass probability 72% with 16 WFA windows - REJECT driven by regime failure and insufficient evidence for deployment.
Verdict: REJECT. Strategy is not viable.

WALK-FORWARD VALIDATION

Time Stability & Overfitting Control
Performance Transfer
In-Sample (IS) vs Out-of-Sample (OOS)

Walk-Forward Analysis - Continuous View

IS (In-Sample) + OOS (Out-of-Sample) equity on a single timeline

Total OOS Return
197877.9%
OOS Win Rate
12 / 16
IS Avg Return
1500.5%
OOS Avg Return
708.4%
Overfitting Score
MEDIUM
Equity Curve - IS vs OOS segments (continuous)
21.917.613.39.14.80.6
ISOOS GoodOOS FragileOOS Failed
OOS performance per period (bar)
P14888.6%1463.0%P21483.4%117.2%P31249.2%1451.2%P4760.0%1019.0%P52096.4%2741.2%P63495.0%359.8%P72145.6%735.6%P83149.4%9.8%P9-1824.4%-1434.0%P101837.4%1939.6%P11-1464.4%-360.8%P121304.0%1372.6%P131311.8%-531.2%P144467.0%-269.2%P15-17.0%1396.4%P16-873.6%1325.0%49%0-49%
ISOOS
WFE (Efficiency):
0.30
Consistency:
82% (9/11)
Performance Degradation:
-53.6%
Failed Windows:4 / 16
Consistency uses only windows with IS > 0. Failed Windows = windows with OOS ≤ 0 or insufficient OOS trades (in some, OOS may still be better than IS). Different denominators.
Overfitting Risk:MEDIUM (n/a)
Professional WFA
Grade:A - ACCEPTABLE
Acceptable for controlled allocation with periodic re-validation.
Pre-verdict module scores (composite; overall verdict and grade above)
WFE Advanced:[?] ROBUST(score 96)(rankWfe 1.368, p 0.927)(pre-verdict composite; overall verdict FAIL)
Overall verdict FAIL - do not rely on this score alone.
Regime: OUTLIER_DETECTED(outliers detected)
Monte Carlo: CONFIDENT(method: Window bootstrap)(P(positive)=99%)[?]
Stress: FRAGILE(recovery: HIGH)[?]
Equity curve: WEAK
Window Breakdown
Period 1
[Fragile]
Opt: 48.9%(84)
Val: 14.6%(32)
█░░░░░░░░░░░30%
Period 2
[Fragile]
Opt: 14.8%(70)
Val: 1.2%(25)
░░░░░░░░░░░░8%
Period 3
[Good]
Opt: 12.5%(86)
Val: 14.5%(36)
███░░░░░░░░░116%
Period 4
[Fragile]
Opt: 15.1%(112)
Val: 10.2%(42)
██░░░░░░░░░░68%
Period 5
[Fragile]
Opt: -5.9%(119)
Val: 27.4%(40)
░░░░░░░░░░░░n/a
Period 6
[Fragile]
Opt: 32.0%(102)
Val: 3.6%(26)
░░░░░░░░░░░░11%
Period 7
[Fragile]
Opt: 18.5%(117)
Val: 7.4%(30)
█░░░░░░░░░░░40%
Period 8
[Fragile]
Opt: 60.9%(138)
Val: 0.1%(40)
░░░░░░░░░░░░0%
Period 9
[Fail]
Opt: -18.2%(96)
Val: -14.3%(20)
░░░░░░░░░░░░n/a
Period 10
[Good]
Opt: 18.4%(70)
Val: 19.4%(26)
███░░░░░░░░░105%
Period 11
[Fail]
Opt: -14.6%(46)
Val: -3.6%(3)
░░░░░░░░░░░░n/a
Period 12
[Good]
Opt: 13.0%(55)
Val: 13.7%(23)
███░░░░░░░░░105%
Period 13
[Fail]
Opt: 13.1%(91)
Val: -5.3%(28)
░░░░░░░░░░░░n/a
Diagnosis: Alpha Reversal (Overfitted)
Period 14
[Fail]
Opt: 44.7%(79)
Val: -2.7%(39)
░░░░░░░░░░░░n/a
Diagnosis: Alpha Reversal (Overfitted)
Period 15
[Fragile]
Opt: -0.2%(73)
Val: 14.0%(35)
░░░░░░░░░░░░n/a
Period 16
[Fragile]
Opt: -8.7%(85)
Val: 13.2%(25)
░░░░░░░░░░░░n/a
Failed Windows Details
Period 9: Validation return is non-positive
Period 11: Validation return is non-positive
Period 13: Validation return is non-positive
Period 14: Validation return is non-positive
▶ Verdict: FAIL
Derived from WFE and consistency thresholds.

PARAMETER SENSITIVITY & STABILITY

Methodology: Sensitivity = R^2 (correlation^2) between parameter value and trial score; we use it as a proxy for 'outcome strongly tied to parameter' (tuning matters). High R^2 = parameter significantly predicts outcome. Magnitude (slope per unit change) is a separate planned metric; Risk Score does not use slope. Sensitivity values: 2 decimal places. Risk Score: integer (floor). From optimization trials or WFA windows.
Suggested Mitigation: Risk Neutral
Parameter
Optimal[?]
Topology[?]
Sensitivity
Status
Adx_trend_threshold
26
0
🟢 Stable
Suggested Mitigation: Risk Neutral
Atr_period
34
0
🟢 Stable
Suggested Mitigation: Risk Neutral
Ema_pullback_threshold
0.01
0.01
🟢 Stable
Suggested Mitigation: Risk Neutral
Hard_stop_atr_mult
4
0.25
🟢 Stable
Suggested Mitigation: Risk Neutral
Loss_time_stop_hrs
74
0.01
🟢 Stable
Suggested Mitigation: Time-decay enforced
Profit_lock
0.02
0.02
🟢 Stable
Suggested Mitigation: Risk Neutral
Profit_stop
0.03
0.04
🟢 Stable
Suggested Mitigation: Risk Neutral
Trailing_gap
0.01
0
🟢 Stable
Suggested Mitigation: Risk Neutral
Trend_exit_adx
16
0.07
🟢 Stable
Suggested Mitigation: Risk Neutral
Scale (classification bands): Scale: round sensitivity to 2 decimals, then band. Stable [0, 0.30); Reliable [0.30, 0.40); Needs Tuning [0.40, 0.60); Fragile >= 0.6. Boundaries: 0.30 = Reliable (start); 0.40 = Needs Tuning (start); 0.60 = Fragile (start). penalisedCount = params with rounded sensitivity >= 0.4. Penalty: 2 per Needs Tuning, 5 per Fragile. Ceiling = 100 - 5xpenalisedCount. Final = max(0, floor(min(Raw, ceiling))). Order: round -> band -> penalisedCount -> Base -> Penalty -> Raw -> Ceiling -> Final. Score: integer (floor).
Sensitivity (R²): strength of linear relationship between parameter value and score (predictability), not magnitude of effect. Slope (impact per unit change) is a separate planned metric; Risk Score uses R² only. From optimization trials or WFA windows.
Topology (when available): curve shape from trials; flat = stable, sharp peak = fragile.
Sensitivity: implemented as R² (correlation²). R² measures strength of linear relationship (predictability), not magnitude of change; we use it as proxy for parameter-outcome tie (high R² = fragility). For true sensitivity (magnitude per unit parameter), derivative-based metric is planned. Values: 2 decimals; Risk Score: integer (floor).
DIAGNOSTIC SUMMARY
1. Local Topology & Stability[?]
Performance Decay (IS ➔ OOS): [?]2.3%
2. Governance Impact (Suggested Mitigation)[?]
Governance metrics below do not affect Risk Score or Deployment; advisory only.
Signal Attenuation: 53.6%
Sharpe Retention (IS ➔ OOS): [?]97.7%
Sharpe Drift (OOS vs IS): [?]-2.3 p.p.
Max Tail-Risk Reduction: [?]21.4%(Risk Reduced)
3. Multi-Parameter Coupling[?]
Coupling analysis: No dominant unstable interactions detected.
AUDIT VERDICT
Deployment Status: REJECTED
Performance Decay: 2.3% (REJECTED if >= 80%).
Final Decision = (Risk Score Verdict) AND (Performance Decay < 80% when Decay is defined; when Decay is N/A this condition is omitted) AND (Min OOS Trades met). REJECTED when any applied condition fails. Performance Decay is a deployment gate (step 2). Governance (Sharpe Drift, Tail-Risk, etc.) is advisory only; does not change the result.
Risk Score: [?]Base 75 − Penalty 0 Status: UNACCEPTABLE (0/100)
Pro-Note: Highest sensitivity: Hard_stop_atr_mult (0.25, Stable).

TRADING INTENSITY & COST DRAG

Execution: Simple (estimated fees)

Results use estimated fees/slippage. Provide exact exchange parameters for Institutional-grade analysis.
Position velocity (holding-period) (3794.8x) is 2.71x institutional turnover (1400.8x). Overlapping positions likely; institutional turnover is used for cost and rebate.
Market Impact (Layer 2.5)
  • ADV $3,845.829 is very low; model assumptions may not hold.
INTERPRETIVE SUMMARY
Net edge remains positive at baseline AUM with manageable execution drag.
Baseline AUM:$500
Avg Trades / Month:27
Annual Turnover (institutional):[?]1400.8x
Position velocity (holding-period)[?]3794.8x
Avg Holding Time:[?]19.8h
Avg position size[?]858.0% of AUM
Cross-check (trades × utilization)[?]~2809.4x
Implied overlap factor[?]2.71x
EFFICIENCY & COST LIMITS
Profit Factor (Gross, full backtest):[?]1.29
Profit Factor (Net, full backtest):1.20
Cost / Edge Ratio:96.4%
Avg Net Profit / Trade (bps)[?]12.27 bps
Break-even Slippage:
Tolerance:n/a (negative edge)
Margin of Safety:Medium
Safety Margin:[?]n/a (negative edge)
BES Status:EDGE DEFICIT
Failure Mode:Spread expansion
COST DECOMPOSITION (CAGR)
Exchange Fees:-561.9%
Slippage:-280.9%
Market Impact (est.)[?]N/A - participation ratio too high for model

Participation ratio exceeds 15% of ADV; square-root model out of range.

Total Cost Drag:-100.0%

Market impact not included (model out of range); total is fees + slippage only.

Rebate Capture:0.48 bps/trade(≈ 6.72% CAGR at current turnover)

Rebate Capture is not included in Total Cost Drag; informational (potential savings with maker-heavy execution).

When gross edge is negative, cost decomposition shows cost allocation; improving execution alone cannot make the strategy profitable.

CAPACITY & MARKET IMPACT
Estimated AUM Capacity:
N/A - model out of range (participation > 15% ADV)
ADV Utilization:
Top 5 traded pairs:n/a
Portfolio weighted:n/a
Market Impact Model:
Assumption:Square-root law
Liquidity regime:Micro / low liquidity
SLIPPAGE SENSITIVITY (NON-LINEAR)
AUM Size
Slippage CAGR
Net CAGR
~$100k[N/A]
N/A
N/A
~$1.0M[N/A]
N/A
N/A
~$5.0M[N/A]
N/A
N/A
~$10.0M[N/A]
N/A
N/A
EXECUTION HARDENING
Order Type Bias:Limit-biased
Taker / Maker Ratio:40 / 60
Limit Fill Probability:59.0%
Opportunity Cost (Fill Decay):17.40 bps
Adverse Selection (Cost):4.00 bps
Latency Sensitivity:Medium
Toxic Flow Risk:Medium
Moderate adverse selection risk in fast markets.
SENSITIVITY TO ALPHA DECAY
Alpha Half-Life:[?]120.0 days
Win Rate Sensitivity:[?]66.0%
RISK & CONTROLS
Primary Constraint:High fee/edge ratio
Gross edge (per trade, at institutional 1400.8x):[?]+72.4 bps

High value reflects low institutional turnover denominator; most capital cost is in overlap periods.

Gross edge (period/CAGR):positive
Available Control Levers:
Reduce trading frequencyLow
Increase entry threshold (signal strength)Low
Shift to maker-only executionLow
STATUS
Deployment Class:Micro-cap / Research-only
COST ADAPTABILITY:❌ FAIL
Required Alpha Boost (bps per trade)[?]0.00 bps
CAPACITY GOVERNANCE:⚠ SCALE-LIMITED
EXECUTION RISK:⚠ WARNING
Confidence Level:High
27 trades/mo, high signal-to-noise
Z-Score: 5.19

STRATEGY ACTION PLAN

Slippage Sensitivity Analysis

Estimated Net Sharpe at each slippage level uses a nonlinear model (power 0.7 in slippage vs reference bps); degradation is not linear in slippage.

When slippage destroys edge, Net Sharpe can go strongly negative. If Sharpe degrades by >30% under 10–15 bps slippage (standard liquidity conditions), the strategy is likely execution-fragile and may not survive live trading. 50 bps stress test shows where the strategy breaks (Net Sharpe < 0).

Slippage (bps)Net SharpeDrawdown Δ (vs baseline)Verdict
0 (Ideal)0.68+0.0%Base Case
5 (Low)0.13200%+ (real: +1353%)🟡 Margin erosion
10 (Avg)-0.43200%+ (real: +2902%)🔴 UNTRADABLE
20 (High)-1.54200%+ (real: +5703%)🔴 UNTRADABLE
50 (Stress)-4.87200%+ (real: +14108%)🔴 UNTRADABLE

Baseline Sharpe: from WFA OOS (window-level).

WFE 0.30 (biased, n=11) / WFE 0.04 (all windows, n=16)

Equity erodes as slippage increases. At 10 bps: Sharpe -0.43, DD +200%. At 50 bps: Sharpe -4.87, DD +200%.

At current pair liquidity, volume limit ~$500. Above that, slippage >10 bps destroys Sharpe. Order-of-magnitude estimate under current assumptions.

The Decision Engine
Phase 1: RE-RESEARCH REQUIRED
  • Governance State: Research Lock (Capital disabled)
  • Allocation: 0% - do not deploy until re-optimized. Add 2 more years of data or reduce parameter count (extend data vs complexity).
  • Monitoring: RE-RESEARCH required (WFE < 0.5, Net Sharpe (10 bps) < 0.2). Check for execution collisions and toxic flow (Adverse Selection).
  • Runtime Kill Switch: Armed (4/16 OOS Fail)
Kill Switch Reset Conditions (ALL must be met):
  • OOS Sharpe > 0 across minimum 2 consecutive windows
  • Fail ratio drops below 33%
  • WFE (all windows) above Phase 2 threshold for this strategy
  • Manual review by risk manager
Phase 2: Full Deployment
  • Trigger: At least 2 consecutive WFA windows with WFE > 0.7 and Trend regime confirmed (Fragile → Stable) (Conservative: Sharpe < 1)

Condition is forward-looking; this report shows one WFE median across all windows. Use WFE (all windows) for this trigger when available; biased WFE (IS > 0 only) is not used for Phase 2.

Why This Works
Bull Case

Theoretical stability only - with real commissions strategy does not survive.

  • Stable stops and volumes protect from black swans
Bear Case (Risks)
  • Only 1/3 regimes pass - not proven across market conditions

Pro-Note: The highest risk is Net Sharpe at 10 bps. Reduce costs or improve edge before scaling.

Recommended Fixes
  • Model Complexity: Simplify logic: reduce indicator count or increase smoothing period. Merge correlated indicators into one signal.(High)
  • Execution: Net Sharpe at 10 bps below 0.2. Reduce costs or improve edge before scaling.(High)

RISK METRICS (OUT-OF-SAMPLE)

Out-of-sample risk metrics from Walk-Forward Analysis (stitched OOS equity curve or window returns).

Max Drawdown[?]36.56%
|
Recovery Factor[?]22.87
Sharpe Ratio (OOS)[?]0.17
|
Sortino Ratio[?]n/a[?]
VaR (95%)[?]-6.72%
|
CVaR (ES)[?]-7.45%
Profit Factor (OOS)[?]1.52
|
Gain-to-Pain[?]0.52
Trade Win Rate[?]n/a
|
Expectancy (loss units)[?]19%(of avg loss)
Period Win Rate (trades)[?]64%
|
Tail Ratio[?]n/a[?]
Payoff Ratio[?]0.85
|
Edge Stability (t)[?]3.78
Skewness[?]-0.00
|
Kurtosis[?]1.11
Durbin-Watson[?]n/a[?]
|
Context: OOS metrics from 1 window (N=470 returns) (small sample - interpret with caution).
Regime Context: High drawdown (Max DD: 36.6%). Consider regime-dependent risk; do not infer volatility expectations without explicit volatility estimate.
Tail Risk Profile: Moderate excess kurtosis (1.11) - heavier tails than normal. Skew: -0.00. Strategy is unprofitable or high drawdown; tail distribution is not the primary concern. ES/VaR ratio: 1.11x. Tail Ratio may be unreliable on small sample.
Tail Authority: Stable tails: losses are tightly clustered around the threshold.
Risk Attribution: Edge driven primarily by hit-rate; payoff profile is modest.
Risk Verdict: Insufficient data - OOS metrics from 1 window are not statistically meaningful. Collect more walk-forward windows before interpreting.
UNSTABLEInsufficient data - single OOS window. Collect more walk-forward windows before interpreting.Max Leverage: 1x

This analysis is for informational purposes only and does not constitute investment advice. Past performance is not indicative of future results. All metrics are model-based and subject to assumptions (slippage, fees, liquidity).