Summary:
1. Replace namespace, table and ds in TableSpec with pvc.HiveDataset
2. Whenever TableSpec is initialized, we initialize a HiveDataset and pass it in so that it's possible to designate multiple partitions.
3. For pvc queries, use table_spec.dataset directly for most of the time.
4. Update all related interfaces to comply with this change.
It turns out TableSpec is defined as a very low level data structure and referenced by many files (>30)... I tried my best to update all of them via code search, unit test and integration test. But to be honest I don't have any context about this rl project at all except knowing this is a fblearner flow pipeline. So please let me know if I miss anything. Thanks!
Reviewed By: kittipatv
Differential Revision: D25538742
fbshipit-source-id: 5df4d36e60d5717c3042a9728daafa86e080a9da
Summary:
Pull Request resolved: https://github.com/facebookresearch/ReAgent/pull/317
Keeping the sanity of OSS API. Merging internal and external classes together means a number of dead options in OSS.
Reviewed By: badrinarayan, kaiwenw
Differential Revision: D23916521
fbshipit-source-id: ad111dabfc1fd354c776a46e765d9ae2c826e292
Summary:
Pull Request resolved: https://github.com/facebookresearch/ReAgent/pull/299
This diff accomplishes several items:
1. Remove PageHandler and consolidate all training functions into one function, using polymorphism to handle model-specific logic
2. Make BatchRunner the sole place where FB vs. OSS context is decided (by choosing FbBatchRunner or OssBatchRunner)
3. Transform ModelManager into a stateless provider.
4. With the exception of model manager, remove all duplicate classes by creating oss & internal versions and using polymorphism, or moving out of workflow/* entirely
5. Replace signals-and-slots API with interfaces
6. Create a DataFetcher class, unifying the APIs to query data on OSS and FB.
Reviewed By: kaiwenw
Differential Revision: D22702504
fbshipit-source-id: 3eb8e93144ca12ac650a4fafc875e29d8ade89e3
Summary:
Pull Request resolved: https://github.com/facebookresearch/ReAgent/pull/294
Can remove them now because automated create_table
Reviewed By: czxttkl
Differential Revision: D22593792
fbshipit-source-id: 81af2bb373a044c03d03deffe6ecf5c16290cff9
Summary:
- The value of state_features are scalars, rather than lists.
- minor changes to query_data
Pull Request resolved: https://github.com/facebookresearch/ReAgent/pull/234
Test Plan: Tox.
Reviewed By: kittipatv
Differential Revision: D21111857
Pulled By: kaiwenw
fbshipit-source-id: 2d690cd3feeb7610e26ae1ffa70add0afc444a77
Summary:
Converted data_fetcher.py’s query_data to PySpark and added sparse2dense logic.
Since we’re using Parquet (which shall be loaded with Petastorm), instead of returning HiveDataSetClass, I’m returning Dataset containing url to parquet.
Pull Request resolved: https://github.com/facebookresearch/ReAgent/pull/223
Test Plan:
Imported from GitHub, without a `Test Plan:` line.
Run `tox`.
Reviewed By: kittipatv
Differential Revision: D21036014
Pulled By: kaiwenw
fbshipit-source-id: 46f1cef7731db365c70dce4831b6d7adca39dce1