In multi-armed bandit problems, KL-UCB is an upper-confidence-bound style algorithm that is asymptotically optimal: its long-run regret matches the best possible bound allowed by the Lai-Robbins lower bound for that problem. This description is adapted from Wikipedia contributors under CC BY-SA 4.0; changes were made. https://creativecommons.org/licenses/by-sa/4.0/
Sources
Wikipedia: Kullback-Leibler Upper Confidence Bound
Reader Challenges (0)
No disputes yet. Spotted an error or a better source? Open the first one.
Sign in to dispute this or suggest a correction.