Krzysztof Duda

pushing every pixel back by how far away it is

reading depth…
scroll ↓

Dots is an iPhone app that takes the camera, a clip or a photo and turns it into a point cloud. there is one dot for every cell of the frame, and each one stands where a depth model says that cell is.

this is a photo of a narrow food alley at night, hung with paper lanterns. scroll, and it comes apart.

the nearest lanterns go first, then the rest of the alley behind them. drag to turn it, or point at a dot to read how far away it is. on a phone, tap it once and tilt the phone to look around.

everything here runs in your browser, on three photos, with the depth Apple's Depth Pro computed for them ahead of time. the rest of the post builds this cloud one step at a time.

a picture already knows its rays

a camera is a pinhole. every pixel looks out from the same point in one direction, and the field of view fixes which direction. pixel (u, v) is a ray:

const x = (u * 2 - 1) * tanX // tanX = tan(horizontal fov / 2)
const y = (1 - v * 2) * tanY
// the point is somewhere along (x, y, 1). depth says where.
const point = [x * depth, y * depth, depth]

so the picture already has every direction. the only thing missing is one number per pixel: how far. give each dot its number and it slides down its own ray.

one row from above is easier to read. the shopfronts lean in on the picture, but here they come out as two straight lines with the alley floor running between them all the way back.

what the model gives back

a depth model reads a single frame and answers with a number for every pixel. here is its answer wiped over the picture, red near and blue far.

the numbers are inverse depth, 1 / metres. close up, a centimetre matters, and 40 m away, ten metres barely do. inverse depth puts most of its range where the detail is, and it is what a pair of eyes measures too: disparity falls off as one over distance.

the rings are even steps of it, each in the colour of its depth. they crowd together near the camera and spread out down the alley.

a lens after the fact

with a distance for every pixel, the photo can be focused again after it was taken. a lens blurs a point by how far it is from the focal plane in inverse depth, so every dot opens into a disc that wide and spreads the same light over all of it. focused on the nearest lantern, the rest of the alley melts into circles. click anywhere on the picture to focus there.

float blur = abs(inverseDepth - inverseFocus); // 0 on the focal plane
float size = cell * (1.0 + 11.0 * blur);
colour /= pow(1.0 + 11.0 * blur, 2.0); // the same light over a wider disc

scroll and the focal plane moves back down the alley. the near lanterns open up and the far ones close. the bright ones turn into the round discs a fast lens makes, because a lantern carries far more light than its pixels can hold, and the blur spreads all of it.

it only knows near from far

Depth Pro gives metres. the model that runs live on the phone, Depth Anything V2 Small on the Neural Engine, doesn't. it knows which pixel is nearer than which, but not by how much in metres. its answer is right only up to two unknowns:

inverse depth = scale · model + shift

halve the scale and the whole alley stretches away from the camera.

a wrong shift is worse. the near end barely moves while the far end folds in towards you, so the straight walls come out bent.

on the phone, ARKit already tracks feature points as the camera moves, and those come with real positions in metres: the white crosses. each frame the model is fitted to them, a weighted least-squares line through (model, 1 / metres).

it is robust: points that disagree a lot get less weight, so a tracker's bad match can't drag the room around. two of these crosses are wrong on purpose, and the fit lets them go.

the same fit works for every source that brings some metres of its own. a Portrait photo carries the depth the phone measured when it was taken, so the photo sets the metres and the model adds the detail. a spatial video has two eyes a baseline apart, and a block matcher on the GPU finds their disparity wherever the match is sure.

dots that land between things

the model runs at a lower resolution than the picture. upsampled, its edges come out soft: a pixel on the fox's outline gets a depth halfway between the fox and the grass behind it. from the front you'd never see it. from the side it is a curtain of dots hanging in the air, here in green.

the fix is to snap. for a dot on a depth edge, look at the neighbourhood around it, split it into the near side and the far side, and send the dot to whichever side's colour is closer to its own. orange fur goes with the fox, and green goes with the grass.

const mid = (nearest + farthest) / 2
// average colour of each side of the edge
const near = mean(colours where depth is nearer than mid)
const far = mean(colours where depth is further than mid)
snapped = distance(colour, near) <= distance(colour, far) ? nearest : farthest

the morph

the picture doesn't turn into dots all at once. the change starts at the nearest thing in the frame and runs back through the depth: the nose goes first, then the muzzle, the eyes, the ears, and the grass last. ahead of the front, the contours are already reading the depth.

in the shader that is one smoothstep per dot. every dot has an order, its depth from 0 at the nearest to 1 at the farthest, plus a little noise so the front breaks up. the front moves past it, and the dot travels to its place along a spring curve that goes a little past and comes back:

float order = pow(t, 0.7) + rand * 0.07;
float p = smoothstep(order, order + band, front);
float q = p - 1.0;
float e = 1.0 + 2.4 * q * q * q + 1.4 * q * q; // a little past it, and back
vec3 pos = mix(card, cloud, e);

the dot also keeps to its pixel's ray the whole way. seen from where the camera stood, every dot is still in front of its own pixel, so the picture never changes while it comes apart. it only shows once you turn.

on the phone

the app does all of this every frame, on the GPU, with TypeGPU on react-native-webgpu: a compute pass that finds the nearest thing and works out every dot, a grid that splits a dot into four where the picture has edges, temporal antialiasing, bloom, and a hologram look that sweeps a sheet of light through the cloud. the frames and the depth reach the GPU zero-copy.

a light that wasn't there

depth also says which way every dot faces: the slope between it and its neighbours. with that, the photo can be lit again by a light that was never there. move the pointer and a lamp follows it into the alley: walls that face it catch the light, and the far end stays dark. on a phone, touch to carry it.

walking in

with true metres and the camera's own lens, the alley is a corridor you can walk down. this is the same cloud seen from where the photo was taken, moving forward along the alley towards its far end. the dots that pass the lens shrink away, and the walls slide by on both sides.

that was an alley on Burano, where every house is painted a different colour. the version on this page is smaller: 97,200 dots, one photo per scene and three.js. the steps are the same. a pinhole gives the rays, a model gives the order, a fit gives the metres, and the edges pick a side.

photos from Unsplash: the lanterns by Caleb Jack, the fox by Patrick van Rooijen, Burano by Yurii Stoian.