Explore a diner and a wharf scene across image, video, audio, 3D, geometry, world navigation, and adapter training. The following examples also cover multimodal retrieval, Earth observation, and model evaluation.
$ mere.run run list --root ./proofs --json⟶ each run stores its manifest, checksums, and outputs
Image → 3D · TRELLIS.2image-3d-trellis2-4b · MLX
Create a 3D asset from one image
SAM 3.1 isolates the wharf fox. MLX TRELLIS.2 converts the cutout into a sealed PBR mesh with shape, texture, and metallic-roughness maps. TripoSR and InstantMesh create lower-detail drafts.
Source plate · generated locallySAM 3.1 cutout · model input
Web viewer: 402k-triangle copy. Full-resolution files remain checksummed in the run manifest.
drag to orbit · pinch or scroll to zoom
Persistent world · Wan 2.2 + DreamXvideo-dreamx-world-5b-ar-mlx
Navigate a persistent scene
One world serve session chains 19 camera moves. Each clip starts from the preceding clip's final frame, which preserves the dock, fox, and fog across camera movement.
Start framesource image
Chunks 1-10pivot right · walk 4.8 m
Chunks 11-13yaw left · 18°
Chunks 14-19forward · 3.6 m
$ mere.run world serve --state-directory ./wharf-state · POST /v1/world/session/transitions {"camera":{"motion":"yawLeft"}}
one session · 19 chained transitionseach chunk continues the terminal state512×288 · 24 fps · M4 Max
MoGe-2 recovers depth, normals, confidence, intrinsics, and a point cloud from the wharf plate. Depth Anything 3 handles multi-view geometry; Video Depth Anything carries depth through motion. Recovered geometry drives the relight.
"The door exhales a draft of ozone and wet asphalt, yielding to a sanctuary of humming neon and scorched lard. Inside, the air holds tobacco smoke and percolating coffee."
— mere.run text chat · 512 tokens$ mere.run text chat --prompt "describe a rain-soaked 1950s diner"
Video · LTX768×512 · 65f · 24fps
Establishing shot, dolly-in
Generated with LTX on Metal from the diner scene description.
Video + Audio · LTX 2.5768×512 · 49f · 24fps
Generate video and audio together
LTX 2.5 converts the diner still into a 2.04-second video with synchronized room tone and rain. Play the MP4 to hear the unamplified AAC track.
$ mere.run video generate "slow dolly forward through the diner" --model video-ltx25-distilled-bf16 --image hero.webp --duration 2 --output-mode audio-video --seed 2121View generation details →
Video → Audio · MMAudiosfx-mmaudio-large-44k-v2
Add synchronized sound to video
MMAudio adds harbor ambience, creaking planks, and gulls to the silent world video. Turn the track on or off to compare the timing.
$ mere.run sfx video generate "rain-damp harbor ambience, slow footsteps creaking on wet wooden planks, water lapping against pilings, distant gull cries, rope tapping a mast" world-walk.mp4 --model sfx-mmaudio-large-44k-v2 --seed 4242
44.1 kHz · synchformer-alignedweights CC-BY-NC · disclosed in model sources
Audio → MIDI · MuScriptormusic-muscriptor-large
Convert a song to MIDI
MuScriptor converts the diner mix into 407 timed MIDI notes across guitars, voice, bass, piano, and drums. Select the original mix or the MIDI render during playback.
$ mere.run music transcribe music_diner.m4a --model music-muscriptor-large -o diner.mid
render: diner.mid → GM soundfontgated weights · CC BY-NC 4.0
Music · ACE-StepG major · 88 BPM · 60s
"Honey, stay one more song with me"
Rockabilly with brushed snare, reverb-heavy Telecaster, doo-wop backing vocals, and tenor sax, generated from one prompt and a lyrics file.
Honey stay one more song with me Underneath the chrome and the canopy Red vinyl shining in the smoky light Save me from the lonely night
Speech · TTS → ASRqwen3-nano · parakeet
Synthesize and transcribe speech
Qwen3 synthesizes the narration, and Parakeet returns a timestamped transcription. Both models run on the device.
[00:00 → 00:08] The door exhales a draft of ozone [00:08 → 00:18] Inside, the air is a thick, amber suspension [00:18 → 00:27] Outside, the rain hammers the plate glass
Music · MiniMax Music 320.000s · 44.1 kHz stereo · seed 2121
Generate a vocal track from lyrics
MiniMax Music 3 generates this 20-second synth-pop track from a style prompt and a lyrics file. The clip contains three supplied lyric lines in 44.1 kHz stereo audio.
$ mere.run music generate "maritime synth-pop with clear English vocals" --model music-minimax-music3 --lyrics-file lyrics.txt --duration 20 --sampling-tier quality --performance-mode optimized --seed 2121
Train · a LoRA from scratchimage-klein-base-9b · rank 16 · 250 steps
Train and compare an image adapter
mere.run generates and captions 24 cyanotypes, validates the training plan, trains the adapter, and renders a held-out fox with and without the adapter. The workflow runs on one MacBook Pro.
M4 Max · 128 GBwall time 2 h 02 m, machine in useadapter 261.2 MB · sha256 starts a1df0d
1 · Dataset generated locally
Review 24 generated cyanotypes
krea2-turbo · seeds 3001-3024captions: vision caption (qwen3-vl)through the resident local API
# start training after you approve the plan
Injected LoRA into 144 FLUX.2 Klein layers.
Training (250/250) loss 0.738672
Saving LoRA artifacts
cyanotype-wharf-r16-250.safetensors# checkpoints at steps 50/100/150/200 saved beside it
3 · Loss curve and checkpoints
Review training loss and checkpoints
step 50step 100step 150step 200step 250
Held-out fox, seed 7777, adjacent checkpoints. The style converges without losing the subject.
4 · Compare the adapter output
Compare the base model and trained adapter
The slider compares one composition with the adapter off and on. Both use seed 7777 and image-to-image strength 0.7; the checkpoint strip shows the text-to-image progression.
Image · Klein LoRA img-to-img13 adapters · scale 1.5
Apply 13 adapters to one image
Klein replays one archival street photo through thirteen private LoRAs. The cyanotype card shows training; the carousel shows finished adapters in use.
64×64×12 source tensor16×16×768 output0.142 s inferenceMetal · local
The color grid shows three principal components of the embeddings. It is not a land-cover class, flood or fire finding, or authoritative geospatial conclusion.
Each result includes the model, runner, plan, artifacts, and result hashes. Scores apply only to Mere's defined test subset and aren't comparable with results from other test sets.
Qwen3.8 27B · low reasoning
71.8%
359 / 500 strict passes · mean 0.8031
Laguna XS 2.1 · built-in runner
49.4%
247 / 500 strict passes · mean 0.6473
Nemotron 3.5 Lightning · built-in runner
34.4%
172 / 500 strict passes · mean 0.5794
Scores use the same 500 text, code, and tool rows. The two text-only models leave 50 vision rows unscored. The evaluation methodology reports Qwen's vision result separately.