Files
Zhengxing Chen cf9a9f2410 Seq2Reward step prediction
Summary:
If we can predict how many steps are remaining at the current step, we can use this information during the planning. We can plan for different look_ahead steps and weight their q-values.

The full context of this diff is as follows. We used to train seq2reward on action sequences of a fixed length and also plan on sequences of the same length. But we find that the SmartAuth dataset, the product we are testing with, has a lot of users with varied, often very short, MDP horizons. So if we always plan on fixed steps ahead, the result would not be good for those users expected to end the MDP much earlier.
The basic idea of this diff is to average over all Q-values under different look_ahead steps, based on the probabilities of how many steps are projected to left for the particular user.
however, I think the current method is still flawed. The step prediction network trains on the data which is the result from the logging policy, not the optimal policy. So I will improve in a future diff for more correct planning.

Reviewed By: kaiwenw

Differential Revision: D25567087

fbshipit-source-id: 9d9d5d9e59e0ce4b1f2e0150e8fa4a1310911cf5
2021-01-16 20:18:52 -08:00
..
2020-08-21 15:59:42 -07:00
2021-01-16 20:18:52 -08:00
2020-08-21 15:59:42 -07:00