A photo comes up. It's either a stock photograph from Unsplash, Pexels, or Pixabay, or it came out of an image model. Swipe right for real, left for generated. I tell you whether you were right, adjust your Elo rating, and serve the next one. The images get harder as you improve and easier as you slip. The categories are people, landscapes, animals, food, architecture, art, and street.
The generated images are hosted-model calls, Flux Schnell on Cloudflare Workers AI, a Hugging Face endpoint, and Gemini 2.5 Flash, five attempts per trigger each. Real photos come in three random categories per trigger. An external cron fires those triggers through the day, so the pool grows in small batches rather than one big nightly run.
Two Elo ratings
Players and images both carry a rating, and every swipe is a match between them. A new player starts at 1200. The expected probability that a player at \(R_p\) correctly classifies an image at \(R_i\) is the usual logistic curve:
\[E = \frac{1}{1 + 10^{(R_i - R_p)/400}}\]A 400-point gap is 10:1 odds, and a 1400 player facing a 1200 image is expected to be right 76% of the time. After each game the player moves by \(K(S - E)\) with \(S\) equal to 1 if they were right, and the image moves the opposite way. Players use \(K = 48\) for their first 30 games and 32 after that. Images always use \(K = 16\), so a single player has half the pull on an image that the image has on them. Both ratings are clamped to 400 and 2400.
I also shrink both ratings 0.5% of the way toward 1200 on every game, so a lopsided pool can't inflate or deflate everyone. The shrink can be larger than a small update, so a correct guess always moves the player up at least one point and a wrong one always moves them down at least one. Without that, a 2300 player who guesses right against an easy image would lose rating.
During a player's first 30 games I skip the image side of the update entirely. Someone testing the app by swiping at random would otherwise drag a batch of image ratings around before their own rating meant anything.
Recalibrating images
Player-vs-image Elo is noisy for images that have only been seen a few dozen times, so a separate trigger recomputes image ratings from the fool rate directly. For each image with 20 or more answers, with \(f\) the fraction of players who got it wrong, the empirical rating is
\[R_{\mathrm{emp}} = \operatorname{clamp}(600 + 1200 f,\; 400,\; 2400)\]So 50% fooled is 1200, 90% fooled is 1680, and 10% fooled is 720. The stored rating becomes \(0.7 R + 0.3 R_{\mathrm{emp}}\), rounded, and is only written if it moved by more than a point. The map is linear rather than the logistic \(1200 + 400 \log_{10}(f/(1-f))\) so it doesn't blow up at \(f = 0\) or \(f = 1\), and with a 0.3 blend the difference is small.
Picking the next image
First I look at whether you've seen more real or more generated images and prefer the one you've seen less. Then I look for approved images you haven't seen within 200 Elo of your rating, ordered by fewest times shown with a random tie-break, and take one at random from the top five. If there aren't any, the window widens to 400, then to anything, and if that still fails the real-or-generated preference is dropped. Within the 200 window your expected score sits between 0.24 and 0.76. A player at 900 sees the images with obvious artifacts. Ordering by times shown means a freshly ingested image gets its 20 answers quickly.
Checking answers on the server
The image endpoint returns an ID, a URL, and dimensions. The is_ai column stays in D1. When you swipe, the server checks that the image was actually served to your device, rejects a second answer for the same image, and throws out responses faster than 300 ms or older than five minutes. There's nothing useful to find in the network tab.
Duplicates in the ingestion pipeline are caught with a SHA-256 of the raw bytes, exact match only. I'd wanted a perceptual hash, but the libraries that compute one need sharp or a canvas, and neither runs in a Worker. The column is still called phash. It catches the same stock photo fetched twice from the same source.
The card is a drag gesture with a 120 px threshold or a 500 px/s flick. A green or red flash lasts 150 ms while the next card, already decoded from a queue of five, slides in underneath. Submission happens in the background, so the swipe never waits on the server.