mirror of
https://github.com/facebookresearch/ReAgent.git
synced 2026-06-16 12:44:41 +00:00
Summary: If we can predict how many steps are remaining at the current step, we can use this information during the planning. We can plan for different look_ahead steps and weight their q-values. The full context of this diff is as follows. We used to train seq2reward on action sequences of a fixed length and also plan on sequences of the same length. But we find that the SmartAuth dataset, the product we are testing with, has a lot of users with varied, often very short, MDP horizons. So if we always plan on fixed steps ahead, the result would not be good for those users expected to end the MDP much earlier. The basic idea of this diff is to average over all Q-values under different look_ahead steps, based on the probabilities of how many steps are projected to left for the particular user. however, I think the current method is still flawed. The step prediction network trains on the data which is the result from the logging policy, not the optimal policy. So I will improve in a future diff for more correct planning. Reviewed By: kaiwenw Differential Revision: D25567087 fbshipit-source-id: 9d9d5d9e59e0ce4b1f2e0150e8fa4a1310911cf5