forge

Execution guide

Forge is a compiler evaluation workspace. Start with an included example, select a candidate recipe and a target, then run a comparison.

Inspect a run

The baseline uses direct MLIR lowering and LLVM O1. The cleanup recipe runs canonicalization and common subexpression elimination before the same lowering, followed by LLVM O2. The aggressive recipe adds loop invariant code motion and LLVM O3.

Each configured MLIR pass executes in a separate tool invocation. Forge verifies the IR after every pass and retains a snapshot. Durations include process startup and isolation overhead. These measurements describe the traced workflow, not an uninterrupted production compiler pipeline.

Use Pass explorer for intermediate MLIR, LLVM diff for optimized IR differences, Assembly for target code, Scheduling for LLVM MCA output, and Provenance for commands and tool versions.

Supported inputs

Inputs are textual MLIR, capped at 32 KiB. The tested examples use func, arith, affine and memref dialects with integer arithmetic and fixed-size memory. Other MLIR that the configured lowering cannot handle produces an explicit compilation failure.

Custom dialect plugins, arbitrary pass flags, remote includes, native code execution, GPU execution and CIRCT are unavailable. Target options are ARM Neoverse N1, ARM Cortex-A78, AMD Zen 3 and Intel Skylake.

Scope

MLIR verification establishes structural well-formedness for each successful stage. Forge does not prove that an optimization preserves program behavior. Operation counts are approximate text statistics used for navigation.

LLVM MCA models scheduling on the selected CPU. Its values are estimates rather than device benchmarks. Whole-function analysis can combine mutually exclusive basic blocks and does not model a workload's full memory behavior. Compare the raw assembly and analysis before drawing performance conclusions.

Workspaces and publication

Guests can run included examples. Create a workspace with a name and password to submit your own MLIR. Runs are private by default. Publishing exposes the source, intermediate IR, assembly, commands and report to anyone with the link. Only publish material you are allowed to share.

Passwords require at least 12 characters. Sessions and API tokens expire after 30 days. Password recovery, enterprise SSO and team invitations are unavailable in this engineering preview.

Bounds

One job executes at a time. The waiting queue holds four jobs. A compiler subprocess has a 30-second wall limit and a 20-second CPU limit; a complete run has a 90-second wall limit. Child processes have memory, output and syscall restrictions. Reports are capped at 8 MiB and workspaces at 100 runs.

Reproduce

Download the report to retain the source, target, recipes, tool versions, command arrays, logs and output hashes. Install the documented LLVM 18 packages and run the repository's compiler CLI. Temporary paths in recorded commands belong to the original run; the CLI regenerates them.

An experiment identity hashes the complete request, tool versions, recipes and runner revision. It identifies an input/configuration combination, not a claim that durations repeat exactly.

Read more

• MLIR pass infrastructure

• LLVM MCA documentation

• API reference

• Privacy and retention