Each heartbeat pushes blood into the capillaries of the face and shifts the skin color by a fraction of a pixel value. A webcam records that shift even though nobody can see it. Remote photoplethysmography (rPPG) is the name for pulling a pulse waveform back out of it. This one runs entirely in the browser from a laptop camera. It estimates heart rate, breathing rate, and two HRV numbers, then paints the estimated pulse back onto the cheeks so there's something to look at.
I've never compared the output against a pulse oximeter or any public rPPG dataset, so I have no error figure, and like every rPPG method it falls apart under motion and changing light.
Finding skin
I use MediaPipe FaceLandmarker for the 478 landmarks, running every third frame because it's the most expensive thing in the loop. From the landmarks I take three boxes: the forehead and both cheeks, each the bounding box of a handful of landmarks padded by 4 pixels. Each box gets downsampled to 32 by 32 and I average R, G, and B over the pixels that pass a skin test. The test is in normalized chrominance, \(r_n = R/(R+G+B)\) and \(g_n = G/(R+G+B)\), and a pixel counts as skin when \(R+G+B > 60\), \(r_n \in (0.35, 0.6)\), and \(g_n \in (0.25, 0.4)\). It throws out background, specular highlights, and shadow. The bounds are hard-coded and untested across skin tones.
The averaging isn't uniform once a heart rate estimate exists. A per-pixel weighting stage keeps 64 frames of green-channel history for the 1,024 forehead pixels. After 30 frames it correlates each pixel's history against a single sine at the current bpm and maps the correlation to a weight \(w_p = 0.1 + 0.9 \max(0, \rho_p)\). Pixels that pulse in time with the estimate count more. The floor of 0.1 keeps any pixel from being silenced completely, so a wrong estimate can still be corrected by the pixels that disagree with it. The reference is a sine only, with no cosine term, so a pixel that pulses a quarter cycle out of phase with the reference scores near zero even though it carries the signal.
POS
The pulse extraction is the Plane Orthogonal to Skin algorithm from Wang et al. (2017). POS works on sliding windows of 48 samples, which is 1.6 seconds at 30 fps. In each window it divides each channel by its window mean to remove the DC level, then projects onto two axes orthogonal to \([1,1,1]\):
\[S_1 = \tilde{G} - \tilde{B}, \qquad S_2 = \tilde{G} + \tilde{B} - 2\tilde{R}\]Green carries most of the pulse. Hemoglobin's strongest absorption peak is actually in the violet, near 415 nm, but that light barely penetrates skin and camera sensors are most sensitive in green, so green wins in practice. Wang et al. then combine the two projections as \(H = S_1 + \alpha S_2\) with \(\alpha = \sigma(S_1)/\sigma(S_2)\). Their argument is that the specular component survives the projection in both signals, and scaling by the ratio of standard deviations makes those leftovers equal in size so they cancel when added. Each window's \(H\) is standardized by \(\sigma(S_1)\) and overlap-added into a 256-sample ring buffer with a \(1/48\) weight. The buffer covers about 8.5 seconds, which turns out to matter later.
Two estimators and a vote
The pulse buffer goes through a 60-sample moving-average detrend and a frequency-domain bandpass to 40 to 200 bpm (zero-pad to a power of two, FFT, zero the out-of-band bins, inverse FFT). Then two estimators run on the same filtered signal. One applies a Hamming window, takes a radix-2 FFT, finds the largest bin in the band, requires it to exceed twice the median in-band magnitude, and refines the position with parabolic interpolation. Its confidence is peak power over total in-band power. The other computes normalized autocorrelation over lags from 0.3 to 1.5 seconds, takes the highest positive lag with the same parabolic refinement, and uses the correlation value itself as confidence.
If the two estimates are within 10% of each other I average them and take the higher confidence. If they disagree the more confident one wins. Anything below 0.08 confidence is dropped. The FFT resolves frequency better on a long record and the autocorrelation is harder to fool with a harmonic, and in practice the disagreement case fires mostly during motion, when neither is right.
Breathing
One breathing source is chest motion. A box below the face, 1.6 face-heights wide and 1.2 tall, is downsampled to 32 by 32 and I take the mean signed green difference between consecutive frames. Signed, because squaring it would double the frequency. This goes through the same detrend, bandpass to 0.1 to 0.6 Hz, and two-estimator vote as the heart rate, with a 900-sample buffer, a 240-sample minimum, and a 0.12 confidence floor.
Another is landmark motion. The mean y-coordinate of the nose tip, nose bridge, and chin oscillates slightly with each breath. I smooth it with a one-pole low-pass at \(\alpha = 0.5\) and buffer it at the 10 Hz detection rate, then run the same estimator pair with a 50-sample detrend and an 80-sample minimum. It tracks head motion at least as much as breathing.
The last is respiratory sinus arrhythmia. Heart rate speeds up on the inhale and slows on the exhale, so the breathing frequency is hiding in the beat-to-beat intervals. I find peaks in the filtered pulse, keep intervals between 0.3 and 1.5 seconds, resample them to a uniform 4 Hz tachogram, and take an FFT of that. The confidence is discounted by 0.8 because it's indirect. It needs 15 peaks, and the pulse buffer is 256 samples, so at 30 fps you need 15 beats in 8.5 seconds, about 98 bpm, and at a resting heart rate the RSA path never fires. When it does fire, the 8.5-second record gives a frequency resolution of roughly 0.12 Hz, about four bins across the whole breathing band, and the 4 Hz resampling doesn't help with that.
One Kalman filter per rate
Heart rate and breathing rate each have a two-state Kalman filter tracking the rate and its derivative under a constant-velocity model. Process noise is 0.25 for heart rate and 0.06 for breathing, because breathing rate should wander more slowly. Every measurement arrives with a variance \(R = R_0 / \max(c, 0.05)\), where \(c\) is the estimator's confidence and \(R_0\) is a hand-picked base variance per source: 4 for the spectral heart rate, 4 for landmark breathing, 20 for chest motion, 12 for RSA. So a confident heart rate reading enters with \(R = 8\) and one at the confidence floor enters with \(R = 80\). An innovation gate drops any measurement more than three sigma from the prediction. The confidence the UI shows is \(1/(1 + P_{00})\), the filter's own state variance, so it only says how steady the estimate is.
When at least 8 valid intervals exist I also compute SDNN and RMSSD with \(n-1\) in the denominator, in milliseconds to one decimal. Both come from camera-detected peaks in a noisy signal, and HRV is very sensitive to peak timing error.
Painting the pulse back on
The display is a single WebGL2 fragment shader with six render targets and ping-pong framebuffers. The visible effect is synthetic, driven by the estimated rate, and no real per-pixel motion gets amplified the way Eulerian video magnification does. Two per-pixel IIR bandpass pairs do run on the GPU for the pulse and breathing bands and their state is kept current, but nothing on screen comes from them. A phase accumulator advances by \(2\pi \cdot \text{bpm} \cdot \Delta t / 60\) each frame using the measured frame interval, so it stays locked to the estimate no matter how the frame rate jitters, and \(\sin\varphi\) drives a flush painted onto two cheek ellipses from an 80 by 60 mask. The flush multiplies the pixel by \((1 + 0.40f, 1 - 0.12f, 1 - 0.14f)\), red up and green and blue down, which is roughly what perfusion looks like. A blend factor ramps toward \(\min(1, 1.5c)\) with a 3-second time constant and decays to zero over a second if the estimate is lost.
Breathing does the same thing with geometry instead of color. A second oscillator at the breathing rate drives a body mask, a trapezoid from the neck to the bottom of the frame, through a vertical smoothstep envelope that anchors the lower edge. The shader lifts the texture coordinate by \(0.05 b\) and narrows the horizontal coordinate about the tracked body center by a factor of \(1 + 0.10 b\). Because the drive is a signed sine, the shoulders rise on the inhale and settle below neutral on the exhale.
Timing and housekeeping
Every bit of signal processing is my own TypeScript on Float64Arrays with no DSP library. The camera runs at 640 by 480 asking for 30 fps, but I never assume 30. I measure the frame interval from video.currentTime, smooth the interval with an EWMA at \(\alpha = 0.05\), and invert it. Smoothing the interval and not the rate avoids the bias from Jensen's inequality. Vitals recompute every 60 frames. There's a 150-frame calibration period before any readout, a 500 ms grace period on a lost face, and a full reset of every buffer and filter if the face is gone for 10 seconds. A shake detector smooths the forehead centroid's frame-to-frame displacement and hides the readouts when it passes 1.2% of the frame.