The previous PRs to make `estree` use codegen broke the swc->estree and
estree->hir conversions. This PR updates those conversions so everything builds
now.
The overall goal of this workstream is to have a Rust representation of ESTree
that we can use as the input and output of the compiler. In Rust environments we
can convert between the native AST of SWC or OXC and ESTree, and when invoked
from JavaScript we can serialize to/from ESTree-compliant JSON. Given that our
first target is to plug into a JS-based compilation toolchain, we need to have a
working serialization to/from ESTree JSON. The point of the codegen-based
`estree` crate is to allow us to model estree as ergonomic, idiomatic Rust (to
make consuming it in code easier) while also allowing us to serialize to/from
spec-compliant ESTree. This PR flushes out one remaining piece.
Updates our estree codegen to generate a custom `Deserialize` implementation
instead of using the derived one from serde. ESTree has some enums whose
variants are themselves enums, for example we have
```rust
enum ModuleItem {
ImportExportDeclaration(ImportExportDeclaration),
Statement(Statement),
}
enum ImportExportDeclaration { ... }
enum Statement { ... }
```
This sort of works with serde's derive implementation: you have to use the "tag"
representation for the inner enums, and an "untagged" representation for the
outer one (ModuleItem). The problem is that with an untagged representation,
serde doesn't know what type of data it's expecting. All it can do is go one by
one and try to parse the data as the first variant (eg ImportExportDeclaration)
then the next one (Statement) and fail when it gets to the end of the list. If
the data isn't valid for any reason, deserialization will fail with a "not a
valid ModuleItem" error. That's true but not helpful, especially if you're
developing estree, are confident that the input json is valid, and need to
figure out where you messed up the definition. It's also not helpful as an
end-user if you're not sure your input json is valid.
So this PR updates our codegen to emit a custom derive implementation that is
identical for both regular enums (like Statement) and recursive ones (like
ModuleItem). We first extract the tag to know what type the value is, then
deserialize exactly as that type. So in the above case, rather than have to
first try parsing every ModuleItem as an ImportExportDeclaration and then fall
through to statement, we just decode the tag (`type` in our case, for example
say it's an "ForStatement"), then deserialize directly as that type (eg, as
ForStatement), then wrap it in the enum variant (ModuleItem::Statement(...)).
For recursive enums like ModuleItem we add an extra wrapper as necessary.
The end result is that we get much more precise errors and deserialization is
more efficient: we always decode just the tag, then as exactly that type.
Note that our serialization is also not perfect right now, because we don't
always emit the `type` key. Serde only emits it when a value appears in an enum.
We can similarly generate custom serializers for all our types to always emit
the tag. That will be straightforward when it's necessary. The current PR was
more of a blocker, because it was really hard to figure out mistakes in the
estree definition given the ambiguous errors. Thanks to this PR we now get
precise errors along the lines of "unknown type `JSXElement`" which are easy to
resolve.
Uses the `syn` crate, which can parse various Rust syntax forms, to parse the
`type` field from json schema description. This allows us to describe complex
types like `"type": "Vec<Option<ArrayElement>>"` directly, rather than requiring
flags like nullable, plural, and nullable_item. The main flag that i'm keeping
is "optional", which is used to indicate when the field itself (not the value)
is optional.
Adds some parts of the ES2015 spec, such as imports and ForOfStatement. This is
enough to get a few more fixtures compiling. The last one uses JSX which I
haven't defined yet.
This is meant to replace the initial `estree` crate with a version that is
generated from a JSON description of ESTree. The idea is to make it easy to
experiment with slightly different representations to balance ergonomic usage of
the data at runtime with serialization compatible with ESTree spec. The JSON
schema looks like this (somewhat abbreviated):
```
{
// Objects are struct types that don't have a `type` and can't appear as an enum
variant
objects: {
Position: {
line: {type: "NonZeroU32"},
column: {type: "u32"}
},
...
},
// Nodes are struct types with a `type` and which can appear as enum variants
(statements, expressions, patterns, etc)
nodes: {
ArrayExpression: {
elements: {
type: "Expression",
plural: true,
nullable_item: true,
}
}
...
},
// Categories of nodes with multiple variants, represented as enums.
// Can be recursive, eg ForInit can be VariableDeclaration or Expression,
// where Expression is also an enum
enums: {
Expression: [
"ArrayExpression",
...
]
},
// Simple enums which have a corresponding string value. Used primarily for
operators (binary/unary/logical/etc)
// but also for things like variable declaration kind (var/const/let)
operators: {
BinaryOperator: {
Plus: "+",
Instanceof: "instanceof",
}
}
}
```
The core estree files are now generated using Cargo's build script mechanism.
Right now i only defined the types and fields from ES5, so i'll have to flush
out the rest of the modern JS spec and extensions like JSX, TypeScript, and
Flow. But already the for-statement example works, showing that this approach
can handle complex cases such as unions of types or other unions (ForStatement
initializer is tricky bc it can be a VariableDeclaration or an Expression - that
works now!).
This is still WIP a bit - now that the ESTree definition is more precise i can
go back and clean up some other code (have to, because the swc -> estree
conversion needs some tweaks now).
Until now i've freely used `panic!`, `unwrap()`, and friends for "error
handling". This PR switches to consistently returning `Result` within the HIR
builder, using a structured error representation that exploits helpers from
`thiserror` and `miette` crates. Miette has a super graphical formatter for
diagnostics as you can see in the screenshot (also see the
[repo](https://docs.rs/miette/5.9.0/miette/index.html)).
This is just a first pass and we'll need to flush out the error handing story
more. Two obvious directions to go next:
* Make HIR construction error-tolerant, so that it can find as many errors as
possible at once rather than failing on the first error. We did this in Relay
Compiler as well, and we can likely borrow some of its helpers.
* Decouple from `miette`. It's very nice but less flexible than I'd like. We can
define our own more generic diagnostic type that contains structured data, then
have a generic conversion mechanism into a miette type so we can use their
display logic.
<img width="789" alt="Screenshot 2023-07-06 at 3 55 38 PM"
src="https://github.com/facebook/react-forget/assets/6425824/e1f1ed4b-5188-4af5-9af4-8f6c5c345023">
Implements support for `ForStatement` from swc -> estree -> hir, flushing out
more of the HIR representation and porting pieces from HIRBuilder as necessary.
Fundamentally this PR is about lowering identifiers during construction of HIR.
For now i'm punting on context variables and assuming all variables are either
global, module-scoped, or locals. The implementation involves a few pieces:
* `estree::Identifer` gets extended with optional binding information. The idea
is that _some_ name resolution mechanism will populate this. Eventually our own,
but we can also borrow data from another source...
* `estree-swc` now configures SWC's (possibly broken?) name resolution mechanism
and sets the above binding data when translating identifiers from swc into our
estree format.
* hir `Builder` tracks identifiers based on `(name, BindingId)` pairs, and
assigns a unique `hir::Identifier` instance for each pair. `Identifiers` are
clone-able (shared).
* Tangential: i updated the printer to handle more instruction variants,
including the now-ported LoadGlobal instr.
I don't love this but it's a start. Long-term we definitely should have our own
name resolution mechanism which we run on the estree prior to lowering to HIR.
Implements a pretty-printer for the HIR and switches the fixture tests to use
this instead of the debug format. It's much more readable now!
Note that not all types are properly printed, I only implemented the
instructions and terminals used in the example. For others we fall back to the
Debug impl so we at least print something.
Adds a new `fixtures` crate intended for running end-to-end tests of the
compiler. As we expand the compiler this will eventually match our JS fixture
setup, where we have .js files as input and produce memoized JS output.
For now, this does the following:
* Parses with SWC (omg this was painful to setup)
* Runs SWC's name resolution, which annotates the SWC ast in-place
* Convert the SWC ast into our `estree` representation
* Convert `estree` into `hir` for each top-level function declaration in the
input program
* Print the Rust `Debug` view of the resulting HIR
As a next step i'll add a pretty-printer for the HIR to roughly match what we
have in JS.
Multiple Place instances can share a reference to a given Identifier in our JS
implementation. For simplicity of the initial port I’m using Rc (for sharing)
and RefCell (for runtime-checked mutability). This is the standard pattern for
shared mutable references in Rust when you don’t need multi-threaded support. We
don’t need HIR to be accessible by multiple threads so this is fine, if we do
multiple threads it will be to parallelize compilation of separate functions.
There are other idioms w less runtime overhead, such as Place holding an index
into a separate vec of identifiers, but that would make the port much less
straightforward.
Starts to port BuildHIR, in the Rust case this means the ESTree -> HIR
conversion. This necessitated flushing out the Builder struct a bit more. Mostly
the logic translates over very directly, and if anything it's cleaner because of
the lack of noise dealing with TypeScript unsoundness for Babel typedefs and
switch statements.
This is a start to porting HIRBuilder, with a largely complete implementation of
`build()`. Notably this includes all the passes which build() calls, and the
helper functions those passes call in turn:
```rust
reverse_postorder_blocks(&mut hir);
remove_unreachable_for_updates(&mut hir);
remove_unreachable_fallthroughs(&mut hir);
remove_unreachable_do_while_statements(&mut hir);
mark_instruction_ids(&mut hir)?;
mark_predecessors(&mut hir);
```
This was pretty straightforward. I ran into one borrow checker issue with (iirc)
`remove_unreachable_for_updates` where i needed to simultaneously hold a mutable
reference into the HIR (to write the ForTerminal) and also get an immutable
reference to check if the update block is reachable. The challenge is this:
```rust
for block in hir.blocks.values_mut() {
^^^^^^^^^^^^^^^^^^^^^^^ hir borrowed mutably here
if let TerminalValue::ForTerminal(terminal) = &mut block.terminal.value {
if let Some(update) = terminal.update {
if !hir.blocks.contains(&update) {
^^^^^^^^^^ borrowed immutably here
terminal.update = None;
^^^^^^^^ mutable borrow still active here (and also bc of the loop)
}
}
}
}
```
I quickly worked around this as we do in the other passes here by first building
a set of the block ids contained in the function (so that the
`hir.blocks.contains()` call becomes a call to the copied set of block ids).
An alternative would be to add a helper function for mutable iteration which
takes the desired value out of the data structure so that you can safely mutate
it and reference the rest of the HIR. Usage might look like this:
```rust
hir.blocks.each_mut(|mut block, hir| {
if let TerminalValue::ForTerminal(terminal) = &mut block.terminal.value {
if let Some(update) = terminal.update {
if !hir.blocks.contains(&update) {
terminal.update = None;
}
}
}
})
```
The lambda would receive the current `block` as a mutable reference, and for the
duration of the call `hir.blocks[block.id]` would be set to a sentinel value
that would crash if accessed. Meanwhile, `hir` would be a readonly reference to
the HIR, allowing the lambda to otherwise lookup information on the HIR but not
mutate it. This seems...kinda fine? But also not immediately necessary as
there's an easy and efficient-enough-workaround for the cases i've encountered
so far.
This is an initial translation of HIR (minus the ReactiveFunction bits). It's
mostly a straightforward translation. A few differences:
* Instead of Effect having an Unknown variant, we type `Place.effect:
Option<Effect>`. Maybe i'll revert that but it seems right.
* Terminal is divided into `Terminal { id: InstructionId, value: TerminalValue
}` and `enum TerminalValue { ... }`, so that we can reference the terminal's id
without caring about which kind of terminal it is.
* Each instruction value variant gets a named struct, whereas in JS we had lots
of anonymous InstructionValue variants.
* Rust generators are still an unstable nightly feature and likely to change
syntax enough that it seems safer not to rely on them. So
`eachTerminalSuccessor()` returns a `Vec<BlockId>`. If the perf of this is bad
enough we can switch to something like SmallVec (a popular crate) which stores a
few elements inline on the stack to optimize for small vecs (which our case
should be).
Perhaps more importantly, what's _not different_: I kept the existing approach
of BasicBlocks having a Vec of owned instruction instances, with
InstructionValue variant operands as Places rather than indices into a shared
instruction array. If this works out it will make porting straightforward.
However there's one aspect of this that is incorrect: right now each `Place`
owns its `Identifier`, whereas in JS the Identifier instances are shared mutable
references. I'll definitely need to refactor the data structures to allow
sharing Identifiers in some form, at which point we may want to change other
things too.
Start of an `estree` crate for representing ESTree with serialization to/from
ESTree JSON. To get the serialization to work quickly I took some shortcuts and
uses a slightly less precise modeling. Longer-term we can write some custom
serialization code and get more precise types (eg avoid the `ExpressionLike`
enum in favor of proper Expression and ExpressionOrFoo types).
lowerIdentifierForAssignment does extra checks such as checking if globals are
lvalues.
UpdateExpression lowers into an assignment and previously we missed out on such
checks.
I incorrectly included these as deps, we were only using these to verify
codegen. It seems fine to leave in but when I imported it internally vscode
would error, so just remove it
- Made most static methods on CompilerError take a single options object as an
argument. With the exception of invariant which takes a condition and an options
object. - Added a new `suggestions` field on CompilerErrorDetail, which we'll
use to provide eslint suggestions - Updated eslint-plugin-react-forget to handle
suggestions
Most of these errors were incorrectly using InvalidInput as a catchall for
rejecting code. I went through each one and manually updated them to be more
accurate
Found this while running Forget on the React tests.
This isn't a high priority because the ESLint plugin would've caught this. But
it'd be nice if either our validation rules caught this or if our compiler did
correctly eliminate the dep array.