Conatus AI

We take AI systems to production. Models are nondeterministic components: we wrap them in deterministic engineering (solvers, validators, release gates, monitoring) and operate the result as a business.

Republic of Korea · founded 2026 · jahn.clawd.monet@gmail.com

What we solve

four situations
  • An AI prototype that must become a reliable, paying product.
  • A document or render pipeline that needs deterministic validation.
  • An inference workload that is too slow or too expensive.
  • Repetitive operations that should run unattended.

Full engagement notes at the end — one address, same-day replies.

00
proof
html+css
11.5 KB gzipped
external requests
0
fonts
system only
analytics
none
built
2026-09-11

Vision & video intelligence

V-JEPA 2.1 · DINOv2-class

Dense self-supervised representations of video — V-JEPA 2.1 and DINOv2-class patch features — turned into working tools that map, search and segment video.

Dense-feature analysis

live · rendered on studio GPUs

Side-by-side loops — source video against its dense-feature PCA-RGB projection, per patch, per frame. One PCA basis is fitted per clip (top-3 components, robust-scaled), so a colour means the same feature direction for the whole clip; the patch grid is upsampled to pixel resolution with an RGB-guided filter and stabilised with motion-adaptive temporal smoothing. Watch the person, sky and ground separate without a single label.

source | PCA-RGB — V-JEPA 2.1 ViT-L/16, 80×45 patch grid; clip generated in-house (Wan2.2)
source | PCA-RGB — V-JEPA 2.1 ViT-L/16, 80×45 patch grid; footage: David Osipov, “Tbilisoba 2024. Georgian dances near Liberty Square” (Wikimedia Commons), CC BY 4.0 — 6 s excerpt, muted
source | PCA-RGB — DINOv2 ViT-g/14, 91×51 patch grid; clip generated in-house (Wan2.2)
source | PCA-RGB — DINOv2 ViT-g/14, 91×51 patch grid; clip generated in-house (Wan2.2)

Embedding atlas

live demo

A video laid out as navigable semantic space: every patch of every frame embedded with V-JEPA 2.1, UMAP-projected to 3D on studio GPUs, one point per patch per frame — 59,840 points from the rainy-plaza clip above, coloured by a PCA-RGB projection of the same patch features. Orbit; hover a point for its patch thumbnail, frame index and timestamp; the dual-thumb scrubber filters by frame range. The map is not an illustration — it is the working index, and the same space powers the moment search below.

59,840 records (1.2 MB) + thumbnail atlas (0.2 MB) — 1.4 MB total, fetched only as this section nears the viewport.

Moment similarity search

live · inside the atlas above

Click any point in the atlas above and the nearest moments across all 59,840 patches light up — an exact cosine ranking computed in your browser over 32-dimensional int8 embeddings, in about two milliseconds, no server round-trip. The query chips run the same highlight from seventeen text prompts whose matches were precomputed offline and hand-checked; three prompts that did not genuinely match the scene were discarded rather than shipped.

Promptable video segmentation

live · rendered on studio GPUs

Open-vocabulary instance segmentation on video: name a concept in plain text — “person”, “skateboard” — and every instance is detected, masked and identity-tracked across the clip, with no fine-tuning. Each identity keeps a fixed colour, so you can watch tracking hold through motion and occlusion: in the street-dance clip, all 73 people in the crowd are tracked individually.

source | instance masks — SAM 3 promptable video segmentation, text concept “person”; all 73 people detected and identity-tracked, one fixed colour per identity. Footage: David Osipov (Wikimedia Commons), CC BY 4.0
source | instance masks — SAM 3, text concepts “person” + “skateboard”; skater and board segmented as separate concepts. Clip generated in-house (Wan2.2)

Generative & inference engineering

kernels · quantisation · gates

Generative and vision models made fast enough — and verified enough — to run as products on our own hardware.

Attention-kernel overlay

production

A custom attention-kernel overlay for an open 14B text-to-video diffusion transformer, running on studio-owned GPUs. Measured on the production workload against the unmodified pipeline:

Overlay results. Reference = unmodified pipeline, identical seeds.
workloadtext-to-video denoising, open 14B diffusion transformer
output736 × 528 px · 81 frames
schedule40 denoising steps
end-to-end denoising time392.9 s → 333.3 s
end-to-end reduction15.2%
mean step time9.06 s → 8.15 s
LPIPS vs reference0.0202
PSNR vs reference34.9 dB
SSIM vs reference0.966
attention backends verified3
kernel fallbacks0
  1. The overlay replaces the attention kernels inside the denoising loop; the rest of the pipeline is untouched.
  2. Query–key scoring runs on an int8 path.
  3. Each of the three attention backends gets its own overlay; none falls back to a stock kernel.
  4. Validation is authoritative; the model is not: a golden-frame comparator checks optimised output against the reference pipeline, frame by frame, under fixed LPIPS / PSNR / SSIM acceptance bands.
  5. The quality gate — LPIPS 0.0202, PSNR 34.9 dB, SSIM 0.966 — held before the overlay was adopted.

Physical AI & simulation

deterministic multi-physics

research prototype · engine output 2026-08-19

A GPU-native, deterministic multi-physics engine for robotic manipulation — custom solver, coupling, validity and determinism architecture, implemented from scratch atop NVIDIA Warp. Rigid-body contact and friction with the static↔kinetic transition, articulated joints and grippers, free-surface liquids (SPH and hybrid grid-particle), liquid-to-surface deposition → film flow → wet-friction coupling, and cloth with self-collision — one GPU pipeline, not a composition of separate simulators.

The coupling is causal: pour a liquid → it deposits on the surface → the film flows → the contact becomes slippery. That chain runs as actual mass transfer inside one CUDA stream, against a closed mass ledger.

Determinism is bitwise. Every subsystem computes in fixed order — no atomics, no hash grids — and full state hashes come back bit-identical across 20 of 20 repeated runs, including with enumeration order reversed. Every step also carries a validity status — SUPPORTED / DEGRADED / INVALID / FAILED — so numerical trouble never hides behind a clamp or a silent fallback: a failure is recorded as a failure.

single NVIDIA RTX PRO 6000 Blackwell-class GPU · Python 3.12 + NVIDIA Warp

Recorded rollouts

Offline renders of actual engine state — the motion is the recorded simulation output, untouched; slow-motion playback (and, where captioned, playback-only frame interpolation) is applied for inspection. Silent H.264 loops.

Dam break — a liquid column (5,070 SPH particles, high-viscosity fluid) collapses in an 8 cm tank, drives a surge wave to the far wall and settles. ~6× slow motion; colour = depth (brighter at the surface) + velocity.
64 × 64 XPBD cloth patch dropped over a table pillar — 240 frames, 4 s simulated, 0.5× speed. Deterministic cloth–rigid contact with no atomics, BVH or CCD; contact penetration 2.6e-8 m in the 32² canonical measurement.
Squeeze and lift — a compliant-contact two-finger gripper squeezes a 0.5 kg box, lifts it 0.12 m and holds. ~4× slow motion; pressure×area vs normal-force self-consistency 1.9e-6, grip success/failure predicted 5/5 across the scene family.

Measured results

Every capability is measured against analytic references and matched open-source engine baselines, under pre-frozen evaluation specs and sealed hold-out scenes; thresholds are never adjusted after seeing results.

Measured on the stated scenes, thresholds frozen beforehand. RTF = simulated time / wall-clock time; 1.0 = real-time.
measurement result conditions
articulation — pendulum period error 0.039% vs analytic; zero energy gain
many-body collision — throughput RTF 4.45 / 4.29 / 3.89 16 / 32 / 64 bodies; 240/240 steps SUPPORTED
liquid — density error P99 ≤ 0.13% all five scenes, validity SUPPORTED
liquid — momentum-ledger residual ~7e-15 dimensionless; ledger closed against external impulses
wet chain — end-to-end mass residual 1.16e-10 kg pour → deposit → film flow → wet contact, one scene; six-stage provenance, three negative controls passed
wet friction — effective static μ vs film thickness 0.428 → 0.320 → 0.212 dry → thin → thick film, monotonic; dry case bitwise-identical to the independent rigid path
cloth — self-collision 0 penetrations · min signed gap +3.7e-5 m vertex–triangle + edge–edge; no atomics, BVH or CCD
compliant grip — success/failure prediction 5/5 squeeze-and-lift scenes; pressure×area vs normal force self-consistent to 1.9e-6
determinism — repeated runs 20/20 bitwise identical all subsystems, full state hash, enumeration-order reversal included

Deterministic solvers

from shipped products

Pure, deterministic solvers from our shipped print products — ported verbatim and running in your browser.

Placement solvers

live demo

First exhibit: the grid packer from a shipped personalised-crossword product, ported to this page verbatim and running in your browser. It is a deterministic, pure constraint solver: it verifies shared-letter connectivity across the answer set, pins the chosen anchor word, then places entries in most-constrained-first order with backtracking, scoring each complete layout and keeping the best.

Edit the word list — ten to fourteen answers — and it re-packs live, reporting attempts, backtracks, elapsed milliseconds and the layout score. Shuffle the anchor to force a different solve.

The same solver packs every paid order at our-story-crossword.

Second exhibit: the seeded word-search engine from the companion print product. Placement is driven entirely by a seeded generator — hash the seed, shuffle candidates, place with bounded backtracking, fill with generator-drawn letters — so the same seed always yields the same puzzle, byte for byte. Determinism here is not a debugging aid; it is the product contract that makes every paid artifact reproducible.

Autonomous operations

the fleet

production

A watcher fleet runs around the clock — six watchers on three-minute to thirty-minute cadences covering marketplace scans, mail, order flow, and bounty and listing deltas — feeding a persistent agent loop that acts under human direction. Synthetic buyer journeys feed it evidence, and when something breaks, the fleet writes the incident note itself.

03:47 bridge session revoked upstream (401) — relogin loop stopped after 2 attempts, ban-risk pattern
03:52 root cause: reply address mis-normalised for anonymised peer ids; fix committed, bridge parked
09:10 re-paired with operator; send path verified end-to-end — incident closed
fleet-written incident note from a live messaging-bridge failure; identifiers sanitised.

3D→text renderer

lab · Glyphit3D — WebGPU, live

lab — exploratory systems, shipped when they beat the incumbent.

Glyphit3D renders live 3D scenes entirely as text: one glyph and two colours per character cell, re-chosen every frame. Selection is driven by the G-buffer rather than luminance alone, and each cell's foreground and background colours solve a continuous-coverage two-colour least-squares fit. It runs in WebGPU at quality Q3 inside a ~50 ms frame budget; without WebGPU the same matcher runs on a CPU worker pool.

Drag the divider to sweep between the native render and the glyph raster, orbit the model, or drop your own .glb/.gltf into the frame.

Left: native 3D render of a sci-fi helmet. Right: the same frame rendered entirely as text — one glyph plus two colors per cell. The un-blur scrubber sweeping across the demo: 3D render on one side of the divider, glyph raster on the other.
Interactive demo needs WebGPU — static capture shown. (Chrome/Edge 113+, Safari 26+.)
Grayscale SSIM vs chafa 1.18.2, both re-rasterised through the same DejaVu Sans Mono atlas; higher is better.
imageours (Q3)chafa (best)
sphere0.98340.9832
torus0.98240.9821
spheres0.98460.9783
DamagedHelmet0.86770.8603
FlightHelmet0.97180.9697
BoomBox0.93010.9251
mean0.95330.9498

Open full-screen — same-origin build: the demo, its worker and its glyph-atlas profiles all load from this site.

Engagement

one address, no form

Good fit:

  • A model demo that must survive paying customers. We wrap it in validators, state machines, release gates and monitoring — then run it.
  • Documents or renders that must come out right every time. We build the pipeline and the gates that prove it.
  • Inference that is the bill or the bottleneck. We do the kernel-level work and verify quality held.
  • Operations that eat a person’s day. We put a watcher fleet and an agent loop on it, under your direction.

Fixed-scope, fixed-price entry points, if you would rather start small than start a conversation:

jahn.clawd.monet@gmail.com

We reply from a running system, usually same-day. One address, no form — there is no backend here to spam. Send the system, the bottleneck, or the failure mode.