FlowBench/
Scientific ML workbenchGitHub ↗
2D Navier–Stokes · held-out evaluation

Inspect a fluid prediction.

Compare a numerical reference with a trained CNN's prediction on held-out test pairs. See exactly where they disagree.

Loading replay data…
Loading…
REPLAY
Relative L2 error
—
This sample only · lower is better
Mean absolute error
—
Stored field units · physical units unknown
Enstrophy error
—
Relative difference in ½·mean(ω²)

Reference, prediction, difference

Reference fieldω reference · numerical solver
0
Target field paired with the input in the dataset (time offset not documented).
PredictionCNN · replay
0
Trained CNN output, computed once and saved.
Absolute error|prediction − reference|
0
Darker regions have larger errors.

Height shows vorticity in the first two plots and absolute error in the third. These are 2D fields drawn in perspective, not 3D fluid simulations.

Grid indices 0–31 · row 0 at the top · physical domain size unknownInput, reference and prediction share one symmetric scale.
Inspect a map to compare a grid cell32 × 32 aligned values per field

Look across a slice

ReferencePrediction
Vorticity in stored units · columns 0–31Dashed map line marks the selected row.

Compare against persistence

This sample only
PredictionRelative L2MAEEnstrophy Δ
Persistence baseline
CNN · replay

Persistence reuses the input field unchanged. These values describe one held-out sample; they are not held-out benchmark performance. The 2000-sample table is in the Benchmark section.

What this workbench checks

THE METHOD
1

Start from a held-out input field

The dataset's own test file (2000 pairs) is never read for statistics, thresholds or training. Validation was carved from the train file by instance ID; the archive exposes no trajectory grouping.

2

Compare with the numerical reference

32 × 32 fields are a stride-4 decimation of the 128 × 128 archive. One mean/std fitted on training pairs normalises inputs and targets; predictions are inverted to stored units before display and metrics.

3

Inspect where it fails

Spatial errors, the persistence baseline and enstrophy are diagnostics, not proofs: enstrophy agreement does not establish a PDE solution, and a 2D surrogate says nothing about 3D Navier–Stokes.

    Held-out benchmark

    2000 test pairs
    All test samples
    ModelParamsRel. L2 meanRel. L2 worstMAEEnstrophy Δ meanEnstrophy Δ worstp50 msp95 ms
    High-vorticity slice
    ModelRel. L2 meanRel. L2 worstMAEEnstrophy Δ mean