[rust] Contributing guide and readmes for every crate

Per the title, this updates the main readme file with a guide to contributing to 
the Rust compiler, and ensures that we have a brief description of every crate 
in local readme.md files. The `forget_hir` one is the most extensive and 
describes the high-level design of the HIR.
This commit is contained in:
Joe Savona
2023-07-14 12:24:19 +09:00
parent 68b0effca3
commit 3e317b8bbf
8 changed files with 143 additions and 2 deletions
+78
View File
@@ -33,3 +33,81 @@ Scaffolding
Reference
- [Babel Plugin Handbook](https://github.com/jamiebuilds/babel-handbook/blob/master/translations/en/plugin-handbook.md)
## Rust Development
## First-Time Setup
1. Install Rust using `rustup`. See the guide at https://www.rust-lang.org/tools/install.
2. Install Visual Studio Code from https://code.visualstudio.com/.
Note to Meta employees: install the stock version from that website, not the pre-installed version.
3. Install the Rust Analyzer VSCode extension through the VSCode marketplace. See instructions at https://rust-analyzer.github.io/manual.html#vs-code.
4. Install `cargo edit` which extends cargo with commands to manage dependencies. See https://github.com/killercup/cargo-edit#installation
5. Install `cargo insta` which extens cargo with a command to manage snapshots. See https://insta.rs/docs/cli/
## Workspace Hygiene
### Adding Dependencies
To add a dependency, add it to the top-level `Cargo.toml`
```
// forget/Cargo.toml
[workspace.dependencies]
...
new_dep = { version = "x.y.z" }
...
```
Then reference it from your crate as follows:
```
// forget/crates/forget_foo/Cargo.toml
[dependencies]
...
new_dep = { workspace = true }
...
```
### Adding new crates
Rust's compilation strategy is largely based on parallelizing at the granularity of crates, so builds can be faster when projects
have more but smaller crates. Where possible it helps to structure crates to minimize dependencies. For example, our various compiler
passes depend on each other in the sense that they often must run in a certain order. However, they often don't need to call each other,
so they can generally be split into crates of similar types of passes, so that those crates can compile in parallel.
As a rule of thumb, add crates at roughly the granularity of our existing top-level folds. If you have some one-off utility code that
doesn't fit neatly in a crate, add it to `forget_utils` rather than add a one-off crate for it.
## Running Tests
Run all tests with the following from the root directory:
```
cargo test
```
The majority of our tests will (should) live in the `forget_fixtures` crate, which is a test-only crate that runs compilation end-to-end with snapshot
tests. To run just these tests use:
```
# quiet version
cargo test -p forget_fixtures
# without suppressing stdout/stderr output
cargo test -p forget_fixtures -- --nocapture
```
Another hint is that VSCode will show a "Run test" option if you hover over a test in the source code, this lets you run a single test easily.
The command line will also give you the CLI command to run just that one test.
## Updating Snapshots
The above tests make frequent use of snapshot tests. If snapshots do not match the tests will fail with a diff, if the new output is correct you
can accept the changes with:
```
cargo insta accept
```
If this command fails, see the note in "first-time setup" about installing `cargo insta`.
@@ -1,3 +1,3 @@
# Build-HIR
This crate converts from ESTree into HIR format as the first phase of compilation.
This crate converts from `forget_estree` into Forget's HIR format as the first phase of compilation.
@@ -0,0 +1,17 @@
# forget_estree
This crate is a Rust representation of the [ESTree format](https://github.com/estree/estree/tree/master) and
popular extenions including JSX and (eventually) Flow and TypeScript.
This crate is intended as the main interchange format with outside code. A typical integration with Forget
will look as follows:
1. Host Compiler parses into the host AST format.
2. Host Compiler converts into `forget_estree`.
3. Host Compiler invokes Forget to compile the input, which (conceptually)
returns the resulting code in `forget_estree` format.
4. Host Compiler convert back from `forget_estree` to its host AST format.
Because Forget is intended to support JavaScript-based toolchains, `forget_estree` is designed to support
accurate serialization to/from estree-compatible JSON. We may also support the Babel AST format
(a variant of ESTree) as well, depending on demand.
@@ -0,0 +1,4 @@
# forget_estree_codegen
This crate is a build dependency for `forget_estree`, and contains codegen logic to produce Rust code to describe the ESTree format
given a JSON schema.
@@ -0,0 +1,4 @@
# forget_estree_swc
This crate converts from SWC's AST format into Forget's `forget_estree` format, which is used as the primary
input/type in Forget's public APIs.
+30 -1
View File
@@ -1,3 +1,32 @@
# HIR
This crate defines the core data structures that Forget uses to represent and compile input programs.
This crate defines the High-level Intermediate Representation (HIR) used by Forget.
While the name is inspired by Rust Compiler's HIR, Forget's HIR is actually quite different.
Rust's HIR is effectively a compact AST, effectively a slightly more canonical form than the
concrete syntax tree produced by the parser.
Forget has two goals that are in tension:
1. Forget needs a detailed understanding of the control-flow semantics and performs sophisticated
data-flow and semantic analysis, all of which benefit from more traditional control-flow graph
representation with flat instruction sequences.
2. At the same time, Forget needs to output code in the original language, and ideally should
produce code that is as similar as possible (for comprehension) and compact (to avoid increasing
bandwidth costs and time to download).
To satisfy both goals, Forget's HIR uses a hybrid of a traditional intermediate representation and
an AST:
1. The HIR is a control-flow graph, with one or more basic blocks each of which contains zero or more
instructions and a terminal node. The blocks are stored in reverse postorder so that compiler passes
can iterate the graph and always visit all predecessor blocks before successors, except in the
presence of loops. This allows many passes to complete in a single pass and eases data flow analysis.
2. Unlike a typical intermediate representation, the HIR uses a rich set of high-level terminal nodes.
Rather than just a list of successors, for example, the terminal type is an enum of variants such as
"if", "for", "for-of", "do-while", and other types to represent expression-level control flow in
JavaScript, such as "ternary", "logical", "optional", "sequence", etc. Notably, these terminals contain
named fields with links to successors (eg "for" has fields for the init, test, update, and body blocks)
but also for the "fallthrough" block, ie the block to the code that comes "after" all the logic of
the terminal. This fallthrough allows Forget to retain the shape of the AST and recover it later in
compilation.
@@ -0,0 +1,3 @@
# forget_optimization
Compiler passes that apply various optimizations to improve the performance and/or size of the program.
@@ -0,0 +1,6 @@
# forget_utils
This is a catch-all crate for utilities and helper code that doesn't have an obvious home elsewhere.
It is expected that this crate will be depended on by lots of other crates in the project. However,
this crate should generally *not* depend on other workspace crates — that's an indication that the
utility you're adding belongs with the crate that uses the utility.