Adaptive learning and bandit VOI
value_of_adaptive_learning_bandit estimates the value of sequential
allocation under UCB, Thompson, or epsilon-greedy policies. It reports
exploration cost, regret, opportunity cost, decision switching, sampling
burden, and stopping diagnostics.
This method is experimental and fixture-backed. Licensed online-allocation data, Rust and retained-binding parity, and mature/stable approval remain external promotion gates.