INTELLIGENCE FOR THE COMPILER

Make every
optimization count.

NeuroCompiler explores a smarter way to optimize programs: combining machine learning and reinforcement learning to guide LLVM’s optimization decisions around measured execution performance.

RESEARCH-DRIVEN · LLVM-BASED · PERFORMANCE-FIRST
OPTIMIZATION SPACE / 001EXPLORING
NC↗
FEATURES
POLICY
LLVM IR
REWARD
STATE → ACTION → MEASURELLVM / ML / RL
01Learn from program structure
02Choose transformations intelligently
03Validate with real execution
THE PROBLEM

One pipeline doesn’t
fit every program.

Traditional optimization levels apply carefully engineered heuristics. They’re powerful—but the best sequence of transformations can vary with a program’s control flow, memory behavior, loops, and target hardware.

⌘
CONVENTIONAL COMPILATION

Fixed strategies

Preset pipelines such as -O2 and -O3 provide strong general-purpose defaults, but may not be ideal for every workload.

◉
THE SUCCESS CRITERION

Measured performance

Smaller intermediate representation does not automatically mean faster code. Execution time is the primary objective, with size and compilation overhead kept in view.

SYSTEM ARCHITECTURE

A learning layer
inside the workflow.

The intended architecture keeps LLVM at the center. NeuroCompiler’s decision layer proposes and evaluates optimization choices; benchmark evidence determines whether those choices help.

01
⌨

C / C++

Source program
+ test inputs

→
02
▤

LLVM IR

Clang frontend
+ feature extraction

→
03
⌘

NeuroCompiler

ML predictor
+ RL policy

DECISION LAYER
→
04
⟳

LLVM passes

Selected valid
transformations

→
05
◫

Benchmark

Correctness, runtime
+ code size

NOTE This is the intended system architecture. Components are being developed and evaluated in phases; the diagram does not imply that every component is already implemented.

RESEARCH ROADMAP

From measurements
to learned decisions.

Reliable optimization begins with trustworthy evidence. The roadmap prioritizes reproducible measurements before model complexity and end-to-end integration.

PHASE 01FOUNDATION

Build the evidence base

Compile diverse benchmark programs under controlled configurations. Verify correctness and collect repeatable runtime, code-size, and compilation-time measurements.

  • Benchmark harness
  • Training-ready dataset
  • Reproducible baselines
PHASE 02LEARNING

Predict, then decide

Train supervised models to estimate optimization outcomes, then investigate an RL policy for choosing useful passes and sequences.

  • Profitability prediction
  • Constrained action space
  • Runtime-aligned reward
PHASE 03VALIDATION

Integrate and evaluate

Connect learned decisions to an LLVM-based workflow and compare results against standard optimization levels and ML-only/RL-only variants.

  • Correctness validation
  • Held-out program testing
  • Transparent performance reports
WHAT MATTERS

Optimize for what
actually runs.

Every experiment should distinguish a promising transformation from a real performance improvement.

T

Execution time

Primary objective. Repeated measurements against a documented baseline.

PRIMARY
B

Code size

Track binary growth or reduction without treating size as a speed proxy.

SECONDARY
C

Compilation overhead

Account for feature extraction, inference, search, and optimization time.

SECONDARY
✓

Correctness & generalization

Reject invalid outputs and test on programs excluded from training.

REQUIRED
OUR ENGINEERING PRINCIPLES
01 /

Measure, don’t assume.

02 /

Keep correctness non-negotiable.

03 /

Report regressions as well as wins.

START A CONVERSATION

Better compiler decisions
start with a question.

Interested in compiler research, benchmarking, collaboration, or learning more about NeuroCompiler? Get in touch.

NEUROCOMPILER / CONTACT
OPEN FOR TECHNICAL DISCUSSION