Skip to content

Writing

Inside TryOnIt: The Architecture of a Client-Side Virtual Try-On SDK

How I structured an open-source virtual try-on SDK as layered TypeScript packages: a DOM-free core, a per-frame fast path, a state machine and renderers.

Published 8 min read

  • Architecture
  • TypeScript
  • MediaPipe
  • WebGL
  • Monorepo

TryOnIt is an open-source SDK that lets shoppers see makeup, glasses, hats, jewellery, watches, hair colour and clothing on themselves through their camera, live on a product page. Everything runs on the shopper’s device: face, hand and body tracking with MediaPipe, and rendering with WebGL2 and three.js.

Three constraints shaped the architecture more than anything else:

  1. Product pages must stay fast. Try-on is an optional feature on a page whose main job is to sell. It cannot cost the page its load time.
  2. Tracking runs 30 times a second. Any design that pushes per-frame data through a UI framework will stutter.
  3. The same products must work on the web, in React and in React Native. A lipstick defined once should look the same everywhere.

This article walks through how those constraints turned into four npm packages, six layers and a few rules that tests enforce.

Four packages, six layers

TryOnIt is a monorepo published as four packages. Each layer may only import from layers with a lower number:

@tryonit/core            pure TypeScript, zero dependencies, no DOM
  Layer 1  domain        asset types, tracking types, errors, landmark indices
  Layer 2  foundations   schema builder, store and state machine, One Euro filter, homography, events
  Layer 3  anchors       face, hand and body anchors, metric scale, asset resolver, requirements

@tryonit/web             browser
  Layer 4  engine        camera, lazy MediaPipe trackers, WebGL2 and three.js renderers, capture, mount()

@tryonit/react
  Layer 5  state         provider and headless hooks
  Layer 6  UI            components, icons, theme, i18n, CSS

@tryonit/react-native    reuses layer 3 and embeds layer 4 inside a WebView

The layering is not just a diagram in a README. ESLint’s no-restricted-imports rule enforces it, so a UI component that reaches into an engine internal, or a core module that imports something browser-only, fails the lint step. Architecture rules that live only in documentation erode one convenient import at a time. Rules that fail a build do not.

A core that does not know it runs in a browser

The most important decision was to keep @tryonit/core completely platform agnostic. It contains the manifest types and validation, the session store and its state machine, the anchor math, the smoothing filters and the asset resolver. None of that needs a browser, so none of it is allowed to touch one:

  • It compiles with lib: ["ES2022"] and no DOM types, so referencing window is a type error, not a code review comment.
  • It never touches window, document or navigator.
  • Network access goes through an injectable fetch, and URLs are resolved without relying on the URL global.
  • The store schedules notifications with queueMicrotask and falls back to a Promise, both of which exist in Hermes, the React Native JavaScript engine.
  • A dedicated test imports the entire public API in a Node environment without DOM globals.

Injecting fetch keeps the resolver usable anywhere, including tooling that runs in CI:

import { resolveAsset } from '@tryonit/core';

const controller = new AbortController();
const asset = await resolveAsset('https://cdn.example.com/aviator.json', {
  signal: controller.signal,
});

// A loader function works too, for merchant APIs.
const fromApi = await resolveAsset((signal) => fetch('/api/tryon/123', { signal }).then((r) => r.json()));

The payoff came when I built the React Native package. Validation, the session state machine, anchors, filters and error codes were reused unchanged, because nothing in them ever assumed a DOM. The full story is in Shipping AR Try-On to React Native and Expo Go.

The per-frame fast path

A try-on session produces a lot of data: hundreds of face landmarks, transformation matrices and segmentation masks, many times per second. If that data flowed through React state, every component subscribed to the session would re-render at camera speed.

So the engine has two separate channels. Per-frame data takes a fast path that never touches the UI layer:

  1. The camera delivers a new frame, and requestVideoFrameCallback wakes the frame loop.
  2. The frame loop hands a downscaled frame and its timestamp to the tracker, which runs MediaPipe at an adaptive 30 or 15 detections per second.
  3. The tracker returns landmarks, a transformation matrix and, where needed, a segmentation mask.
  4. The tracking pipeline smooths the results with One Euro filters and computes anchors and the metric scale.
  5. On every requestAnimationFrame, the renderer draws the latest frame state: it uploads the frame texture and masks, then composites the effect or renders the 3D scene.

The UI store only hears about things a person would notice: the face became visible or was lost, the status changed, and performance numbers, which update at most twice per second.

Three rules fall out of this design:

  • Per-frame data never enters the store. Landmarks, anchors and masks stay in the pipeline, so React re-renders stay rare.
  • Render every frame, detect when needed. Detection runs once per new camera frame at an adaptive rate. Rendering reuses the latest smoothed pose on every animation frame, so motion stays fluid even when detection slows down.
  • Losing tracking is graceful. When the face disappears, the last pose is held for 300 ms while the product fades out. Only then does tracking.faceVisible flip to false, which shows the “look at the camera” hint. A blink of lost tracking never flashes the UI.

Modelling the session as a state machine

A camera session has more asynchronous edges than it first appears: a permission prompt, model downloads, the tab being hidden, a product switch that suddenly needs a different tracker, errors at any stage and a retry button. Expressing that with a handful of booleans invites impossible combinations, such as “loading models” and “paused” at the same time.

TryOnIt models the session as a finite state machine instead:

From Trigger To
idle start() checking
checking live camera session requesting-camera
checking startFromImage() loading-models
requesting-camera camera stream available loading-models
loading-models models loaded ready
ready tracking starts running
running pause() or tab hidden paused
paused resume() running
running setAsset() needs a new tracker loading-models
checking, requesting-camera, loading-models failure error
error retry checking
running stop() idle
idle, running destroy() destroyed

Invalid transitions are ignored and logged in debug mode, so a stray resume() after destroy() cannot resurrect a dead session. The complete transition table lives in one constant, and a unit test checks every pair of states against it.

The store itself is deliberately small: immutable state, notifications batched per microtask, and an API compatible with React’s useSyncExternalStore. A subscribeSelector helper mirrors slices into Redux, Zustand or anything else, which is how the React hooks subscribe to exactly the slice a component needs.

How each product type is drawn

Different products need different rendering techniques, so the engine has a renderer per family:

  • Makeup. For each layer, region polygons built from face landmarks are rasterised into a half-resolution mask, feathered with a separable Gaussian blur and composited over the frame. The composite modulates the product colour by the brightness of the pixels underneath, so skin and lip texture survive instead of being painted over. Lips subtract the inner mouth polygon, so teeth stay white when the shopper smiles.
  • Hair colour. A segmentation mask drives a recolour that keeps the original luminance pattern, so strands and shine stay visible.
  • 2D stickers. Face stickers are quads anchored to face points and rotated with head roll.
  • Clothing. A garment image is drawn as an 8 by 8 grid and warped with a homography from four garment anchor points onto the detected torso.
  • 3D products. Models are authored in millimetres and placed in centimetre camera space by a perspective camera that matches MediaPipe’s face geometry (a 63 degree vertical field of view). Invisible, depth-only occluders, a head ellipsoid and wrist and finger cylinders, hide the parts that should be behind the body, such as glasses temples.

Mirroring is a single CSS transform on the stage wrapper. Video, landmarks, the WebGL canvas and the three.js layer always agree because they are mirrored together, and photo capture applies the same transform so the saved picture matches the preview.

The 3D pipeline only works if models follow strict unit and axis conventions. I cover those in Preparing 3D Models for AR Try-On.

Loading only what a product needs

Every asset type maps to the trackers and renderers it requires through getAssetRequirements(asset). A lipstick needs the face tracker and the makeup renderer; a watch needs the hand tracker and three.js. The engine loads those chunks on first use and keeps them warm until destroy(), so switching between two lipsticks downloads nothing new. This requirements map is the backbone of TryOnIt’s performance story, which I wrote up separately in Keeping Web AR Fast.

Adding a new asset type in six steps

A good test of an architecture is how much of the system a new feature touches, and whether the tooling tells you what you forgot. Adding a product type to TryOnIt follows a fixed path through the layers:

  1. Types. Add the props interface and asset alias in the domain layer and extend the AssetType and AssetManifest unions.
  2. Validation. Add a schema shape with defaults and register it. The compiler checks that the schema’s output type matches AssetManifest, so the two cannot drift.
  3. JSON Schema. Add a branch to the published schema. A sync test fails until required fields, enums and defaults match the runtime validator.
  4. Requirements. Map the type to its trackers and renderer.
  5. Anchors and rendering. Add pure, fixture-tested anchor math in core, and handle the type in a renderer (or add a new lazily imported one).
  6. Docs and samples. Document the fields, add a sample asset, and add UI hints if the type introduces a new tracker.

Steps 2 and 3 are the interesting ones: the type system and a test make forgetting part of the work impossible to ship.

Takeaways

If I had to compress the architecture into a few principles that apply beyond try-on:

  • Put platform-independent logic in a package that cannot import the platform. Type settings and a no-DOM test turn a good intention into a guarantee.
  • Separate the hot path from UI state. Let the UI subscribe to the few changes a person can perceive, not to every frame.
  • Model lifecycles with explicit states. A transition table is easier to test, and to reason about, than a set of flags.
  • Let tooling enforce the rules. Lint for layer boundaries, the compiler for schema drift, tests for invariants.

All four packages are MIT licensed on npm: @tryonit/core, @tryonit/web, @tryonit/react and @tryonit/react-native.

More articles

Have an app in mind?

Tell me what you're building. I reply to every serious enquiry within two working days.