The showcase

Explore local runs
from start to finish

Explore a diner and a wharf scene across image, video, audio, 3D, geometry, world navigation, and adapter training. The following examples also cover multimodal retrieval, Earth observation, and model evaluation.

$ mere.run run list --root ./proofs --json each run stores its manifest, checksums, and outputs
Image → 3D · TRELLIS.2 image-3d-trellis2-4b · MLX

Create a 3D asset from one image

SAM 3.1 isolates the wharf fox. MLX TRELLIS.2 converts the cutout into a sealed PBR mesh with shape, texture, and metallic-roughness maps. TripoSR and InstantMesh create lower-detail drafts.

Generated source plate: a red fox in a yellow raincoat on a foggy wharf.
Source plate · generated locally
The fox cut out from the plate with SAM 3.1, transparent background.
SAM 3.1 cutout · model input
$ mere.run vision image-to-3d-trellis2 fox-cutout.png --seed 43 -o ./fox-mesh
.glb textured, baked atlas206.1 MB · sha256 starts b7cfc5
.glb vertex-color geometry62.6 MB · sha256 starts 979ea6
.obj + .ply + .pbrvox PBR voxels236 + 60.4 + 36.8 MB
2,236,756 triangles 1,021,504 PBR voxels seed 43 512³ O-Voxel pipeline
Web viewer: 402k-triangle copy. Full-resolution files remain checksummed in the run manifest.
Two rendered views of the reconstructed textured fox mesh. drag to orbit · pinch or scroll to zoom
Persistent world · Wan 2.2 + DreamX video-dreamx-world-5b-ar-mlx

Navigate a persistent scene

One world serve session chains 19 camera moves. Each clip starts from the preceding clip's final frame, which preserves the dock, fox, and fog across camera movement.

Starting frame of the world session: the generated fox wharf plate. Start framesource image
Chunks 1-10pivot right · walk 4.8 m
Chunks 11-13yaw left · 18°
Chunks 14-19forward · 3.6 m
$ mere.run world serve --state-directory ./wharf-state · POST /v1/world/session/transitions {"camera":{"motion":"yawLeft"}}
one session · 19 chained transitions each chunk continues the terminal state 512×288 · 24 fps · M4 Max
Geometry and VFX · source image → assets MoGe-2 · DA3 · VDA

Extract geometry from one image

MoGe-2 recovers depth, normals, confidence, intrinsics, and a point cloud from the wharf plate. Depth Anything 3 handles multi-view geometry; Video Depth Anything carries depth through motion. Recovered geometry drives the relight.

metric depth, in metres EXR + PLY + camera JSON
The original generated wharf plate with the fox in a yellow raincoat.
1 · platekrea2 + klein
Metric depth map of the wharf plate.
2 · metric depthmoge2 · .exr
Surface normal map of the wharf plate.
3 · normalsmoge2 · .exr
Per-pixel validity and confidence mask of the recovered geometry.
4 · confidencevalidity mask
Recovered 3D point cloud of the wharf scene viewed from off-axis.
5 · point cloud.ply
Reconstructed textured fox mesh, two views.
6 · meshtrellis2
The wharf plate relit from foggy day to night lamplight.
7 · relightday → night
Portrait with detected body, hand, and face landmarks drawn over it.
Pose

Detect 107 landmarks

Body and face points with confidence scores, ready for compositing.

$ mere.run vision pose portrait.png --json
Dense optical flow field visualized as a color wheel image.
Optical flow

Measure optical flow

The diner dolly-in contains 393,216 optical-flow vectors in a Middlebury .flo file. Mean motion is 48.4 px.

$ mere.run vision flow a.png b.png -o shot.flo --accuracy very-high
Image + Vision · klein-nano + Falcon Perception
Generated 1950s diner scene with Falcon Perception object detection and segmentation overlays.
waitress · 0.95 jukebox · 0.92 neon sign ×3
Text · Creative gemma4 · 0.85t
"The door exhales a draft of ozone and wet asphalt, yielding to a sanctuary of humming neon and scorched lard. Inside, the air holds tobacco smoke and percolating coffee."
— mere.run text chat · 512 tokens $ mere.run text chat --prompt "describe a rain-soaked 1950s diner"
Video · LTX 768×512 · 65f · 24fps

Establishing shot, dolly-in

Generated with LTX on Metal from the diner scene description.

Video + Audio · LTX 2.5 768×512 · 49f · 24fps

Generate video and audio together

LTX 2.5 converts the diner still into a 2.04-second video with synchronized room tone and rain. Play the MP4 to hear the unamplified AAC track.

$ mere.run video generate "slow dolly forward through the diner" --model video-ltx25-distilled-bf16 --image hero.webp --duration 2 --output-mode audio-video --seed 2121 View generation details →
Video → Audio · MMAudio sfx-mmaudio-large-44k-v2

Add synchronized sound to video

MMAudio adds harbor ambience, creaking planks, and gulls to the silent world video. Turn the track on or off to compare the timing.

$ mere.run sfx video generate "rain-damp harbor ambience, slow footsteps creaking on wet wooden planks, water lapping against pilings, distant gull cries, rope tapping a mast" world-walk.mp4 --model sfx-mmaudio-large-44k-v2 --seed 4242
44.1 kHz · synchformer-aligned weights CC-BY-NC · disclosed in model sources
Audio → MIDI · MuScriptor music-muscriptor-large

Convert a song to MIDI

MuScriptor converts the diner mix into 407 timed MIDI notes across guitars, voice, bass, piano, and drums. Select the original mix or the MIDI render during playback.

0s6s12s18s24s
guitars voice bass piano drums 407 notes · 6 instruments
$ mere.run music transcribe music_diner.m4a --model music-muscriptor-large -o diner.mid
render: diner.mid → GM soundfont gated weights · CC BY-NC 4.0
Music · ACE-Step G major · 88 BPM · 60s

"Honey, stay one more song with me"

Rockabilly with brushed snare, reverb-heavy Telecaster, doo-wop backing vocals, and tenor sax, generated from one prompt and a lyrics file.

Honey stay one more song with me
Underneath the chrome and the canopy
Red vinyl shining in the smoky light
Save me from the lonely night
Speech · TTS → ASR qwen3-nano · parakeet

Synthesize and transcribe speech

Qwen3 synthesizes the narration, and Parakeet returns a timestamped transcription. Both models run on the device.

[00:00 → 00:08] The door exhales a draft of ozone
[00:08 → 00:18] Inside, the air is a thick, amber suspension
[00:18 → 00:27] Outside, the rain hammers the plate glass
Music · MiniMax Music 3 20.000s · 44.1 kHz stereo · seed 2121

Generate a vocal track from lyrics

MiniMax Music 3 generates this 20-second synth-pop track from a style prompt and a lyrics file. The clip contains three supplied lyric lines in 44.1 kHz stereo audio.

$ mere.run music generate "maritime synth-pop with clear English vocals" --model music-minimax-music3 --lyrics-file lyrics.txt --duration 20 --sampling-tier quality --performance-mode optimized --seed 2121
500 semantic frames 30 flow steps BF16 optimized 2 speech transcripts
Lyrics in the published clip
Harbor lights are drawing silver lines
Night rolls in, the signal still shines

Keep the whole world close tonight
View generation details →
MiniMax Music 3 Community License · model revision included
SFX · Woosh DFlow 3.5s · 48 kHz · seed 1963

Generate a ceramic impact sound

Woosh generated a ceramic mug shattering on diner tile.

$ mere.run sfx generate "single ceramic coffee mug dropped onto hard diner tile, sharp impact and scattered fragments" --model sfx-woosh-dflow --duration 3.5 --cfg 4.5 --seed 1963
Code · Swift qwen3-coder · streamed
mandelbrot.swift
/// Computes the Mandelbrot set for a given grid of complex points.
func generateMandelbrotSet(
    width: Int,
    height: Int,
    bounds: ComplexPlaneBounds = .default,
    maxIterations: Int = 100
) -> [[Int]] {
    var mandelbrot: [[Int]] = .init(repeating: .init(repeating: 0, count: width), count: height)
    let xStep = (bounds.right - bounds.left) / Double(width - 1)
    let yStep = (bounds.bottom - bounds.top) / Double(height - 1)
    for y in 0..<height {
        for x in 0..<width {
            let cx = bounds.left + Double(x) * xStep
            let cy = bounds.top + Double(y) * yStep
            mandelbrot[y][x] = iterate(cx: cx, cy: cy, max: maxIterations)
        }
    }
    return mandelbrot
}
compiled · rendered · 0.30s
Mandelbrot overview rendered locally
Seahorse Valley zoom rendered locally
Vision · Falcon Perception grounding + masks

Ground objects with text prompts

Falcon Perception detects prompted objects in the diner image, and SAM 3.1 creates their masks.

waitressbox (0.61, 0.28) → (0.74, 0.68) jukeboxbox (0.84, 0.28) → (0.96, 0.65) neon sign3 detections · masks saved
$ mere.run vision ground hero.png --query "waitress"
OCR · Ideogram JSON → LightOn 1024×1024 · seed 1964

Generate and read structured text

Ideogram generates the ticket from structured text, and LightOn OCR extracts each line.

MERE DINER
TABLE SEVEN
_________________________
COFFEE
CHERRY PIE
_________________________
ORDER READY
$ mere.run vision ocr ocr-receipt-ideogram.png --backend lighton --quiet
Ideogram generated diner order ticket reading MERE DINER, TABLE SEVEN, COFFEE, CHERRY PIE, ORDER READY.
Train · a LoRA from scratch image-klein-base-9b · rank 16 · 250 steps

Train and compare an image adapter

mere.run generates and captions 24 cyanotypes, validates the training plan, trains the adapter, and renders a held-out fox with and without the adapter. The workflow runs on one MacBook Pro.

M4 Max · 128 GB wall time 2 h 02 m, machine in use adapter 261.2 MB · sha256 starts a1df0d
1 · Dataset generated locally

Review 24 generated cyanotypes

Contact sheet of the 24 cyanotype training images: maritime subjects printed white on Prussian blue.
krea2-turbo · seeds 3001-3024 captions: vision caption (qwen3-vl) through the resident local API
2 · Validate the training plan

Check the plan before training

$ mere.run image train-lora --data ./dataset \ --model image-klein-base-9b --rank 16 \ --training-steps 250 --batch-size 2 \ --checkpoint-interval 50 --seed 2121 \ --gradient-checkpointing \ --output cyanotype-wharf-r16-250.safetensors \ --preflight --json { "mere_run_version": "0.21.0", "status": "ok", "summary": "24 usable pair(s), ready to train.", "result": { "dataset": { "usable_pair_count": 24, "missing_caption_count": 0, "duplicate_caption_count": 0 }, "model": { "requested": "image-klein-base-9b", "installed": true }, "plan": { "rank": 16, "training_steps": 250, "learning_rate": 0.0005, "max_resolution": 512, "checkpoint_interval": 50, "expected_checkpoint_count": 5 } } }
# start training after you approve the plan Injected LoRA into 144 FLUX.2 Klein layers. Training (250/250) loss 0.738672 Saving LoRA artifacts cyanotype-wharf-r16-250.safetensors # checkpoints at steps 50/100/150/200 saved beside it
3 · Loss curve and checkpoints

Review training loss and checkpoints

step 10step 130step 2500.880.680.49
Held-out fox prompt rendered with the step-50 checkpoint.
step 50
Held-out fox prompt rendered with the step-100 checkpoint.
step 100
Held-out fox prompt rendered with the step-150 checkpoint.
step 150
Held-out fox prompt rendered with the step-200 checkpoint.
step 200
Held-out fox prompt rendered with the final step-250 adapter.
step 250

Held-out fox, seed 7777, adjacent checkpoints. The style converges without losing the subject.

4 · Compare the adapter output

Compare the base model and trained adapter

The slider compares one composition with the adapter off and on. Both use seed 7777 and image-to-image strength 0.7; the checkpoint strip shows the text-to-image progression.

Base Klein 9B render of a fox in a raincoat on a wharf: photographic.
The same wharf image restyled through the trained cyanotype adapter: paper border, Prussian wash, identical composition.
base klein-9b + cyanotype lora
$ mere.run image generate --model image-klein-9b --seed 7777 -i base.png --strength 0.7 --lora ./cyanotype-wharf-r16-250.safetensors
prompt held out of dataset img2img restyle · strength 0.7 run inspect verified
Image · Krea 2 LoRA 1280×720 · 16 steps

Compare a base model and an adapter

Krea 2 Turbo renders the diner prompt twice: base model, then custom LoRA at scale 2.0. Material, face shape, and miniature texture change.

$ mere.run image generate --model image-krea2-turbo --steps 16 --lora ./custom-style.safetensors --lora-scale 2.0
Seed 7777 · unedited model output
Krea 2 Turbo baseline showing a human night manager counting coins in a rainy 1950s diner booth.
Base Krea
Krea 2 Turbo with a custom LoRA applied, shifting the diner booth scene toward stop-motion puppet texture.
Custom LoRA
Image · Klein LoRA img-to-img 13 adapters · scale 1.5

Apply 13 adapters to one image

Klein replays one archival street photo through thirteen private LoRAs. The cyanotype card shows training; the carousel shows finished adapters in use.

$ mere.run image generate --model image-klein-9b --ref-image street-kiss.png --strength 0.55 --lora ./blade-runner-rain.safetensors --lora-scale 1.5
Local LoRA image-to-image outputs · no semantic post-editing
Klein 9B generated diner scene using a Krea diner reference and a dog reference, showing the dog seated alone in a rainy 1950s booth.
Image · Klein 9B references 1280×720 · 12 steps

Combine scene and subject references

Klein 9B combines the Krea diner as an environment reference with a generated dog portrait as the subject, placing the dog in the booth.

$ mere.run image generate --model image-klein-9b --ref-image diner.png --ref-image dog.png --prompt "swap the dog into the diner"
First reference image, baseline Krea diner booth render.
Reference A · Diner
Second reference image, generated brown and white dog portrait.
Reference B · Dog
Geo · Sentinel-2 → OlmoEarth Halifax · September 29, 2025 · MLX

Create spatial features from satellite data

OlmoEarth v1.2 Base converts a public 12-band Sentinel-2 L2A observation over central Halifax into a 16×16 grid of 768-dimensional features.

$ mere.run geo olmoearth halifax.safetensors --model vision-embed-olmoearth-v12-base --patch-size 4 --input-resolution 10 -o embeddings.safetensors
64×64×12 source tensor 16×16×768 output 0.142 s inference Metal · local
The color grid shows three principal components of the embeddings. It is not a land-cover class, flood or fire finding, or authoritative geospatial conclusion.
True-color preview of a public Sentinel-2 L2A tile over central Halifax acquired September 29, 2025.
Sentinel-2 L2A · true color
PCA color visualization of the 16 by 16 OlmoEarth spatial embedding grid.
OlmoEarth · 768d features → PCA color
Multimodal retrieval · Qwen3-VL 2B 256d Matryoshka · cosine similarity

Retrieve images with text

Qwen3-VL embeds one text query and three showcase images in the same normalized vector space. The matching fox image ranks first.

Text querya red fox in a yellow raincoat standing on a foggy wharf
Top-ranked result: a red fox in a yellow raincoat on a foggy wharf.
#1 · 0.799
Second-ranked result: the neon diner scene.
#2 · 0.268
Third-ranked result: a Mandelbrot seahorse fractal.
#3 · 0.087
Query ↔ fox
0.799

The matching subject, clothing, and setting rank first.

Query ↔ diner
0.268

A cinematic scene, but the wrong subject and place.

Query ↔ fractal
0.087

The unrelated abstract image ranks last.

$ mere.run vision embed --text "a red fox in a yellow raincoat" --image fox.webp --image diner.webp --image fractal.webp --dimensions 256
Evaluation · reproducible results Mere comprehensive subset · 500 comparable non-vision rows

Compare reproducible evaluation results

Each result includes the model, runner, plan, artifacts, and result hashes. Scores apply only to Mere's defined test subset and aren't comparable with results from other test sets.

Qwen3.8 27B · low reasoning
71.8%
359 / 500 strict passes · mean 0.8031
Laguna XS 2.1 · built-in runner
49.4%
247 / 500 strict passes · mean 0.6473
Nemotron 3.5 Lightning · built-in runner
34.4%
172 / 500 strict passes · mean 0.5794
Scores use the same 500 text, code, and tool rows. The two text-only models leave 50 vision rows unscored. The evaluation methodology reports Qwen's vision result separately.
$ mere.run eval promote ./report.json --output ./receipt.json
Each run stores its command, manifest, checksums, and outputs on disk. Evaluation results include immutable result hashes.