Entropy reduction is always bounded by a deterministic policy’s entropy reduction.
In particular, entropy reduction for a blind policy is a convex function on the convex set ΔA of blind policies; so an extreme point will always be a maximum.
Likewise, we can get a weaker form of the bound: ΔH ≤ ΔH_f ≤ Δ H^{max}{blind} + H_f (A)
where f is an optimal (or at least better than the actual policy) sighted deterministic policy, and H_f means we take the entropy under the distribution using f as the policy.
This doesn’t seem to make the proofs easier, but it’s an interesting fact anyways
Entropy reduction is always bounded by a deterministic policy’s entropy reduction.
In particular, entropy reduction for a blind policy is a convex function on the convex set ΔA of blind policies; so an extreme point will always be a maximum.
Likewise, we can get a weaker form of the bound: ΔH ≤ ΔH_f ≤ Δ H^{max}{blind} + H_f (A)
where f is an optimal (or at least better than the actual policy) sighted deterministic policy, and H_f means we take the entropy under the distribution using f as the policy.
This doesn’t seem to make the proofs easier, but it’s an interesting fact anyways