Tools
The machine on your desk.
Everything here runs where you are, a dozen models and a handful of classical algorithms, downloaded once and computed in this tab. The files you drop are sent nowhere, because nothing needs to be.
Runs entirely in your browser. Nothing is uploaded.
Vision
Language & text
Audio
The models behind these
These tools are other people’s research. Every model below is downloaded into your browser and run there, here is who built it, what it does, and the licence it carries.
Reads printed text out of images and scans, line by line, using a recurrent network rather than character templates.
Names what a photo shows, from a thousand ImageNet categories: a vision model small enough for a phone.
Finds individual objects in an image and returns a box for each, rather than labelling the picture as a whole.
Estimates 17 body points of a single person (shoulders, elbows, knees) and so their posture.
Transcribes spoken audio, detecting the language for itself as it goes.
Translates between German and English; a separate small model for each direction.
Finds names, places and organisations in text: the basis of the redaction tool.
Judges whether an English sentence reads as positive or negative.
Answers a question by marking the span of the supplied text where the answer sits.
Condenses longer text into a few sentences.
Sorts text into categories you type yourself, having never been trained on them.