Barcode Detection and Shape Detection API: Browser-Native Computer Vision Internals
Real-World Problem Context
Web applications increasingly need to scan barcodes for inventory management, detect faces for photo organization or AR filters, and recognize text in images for document processing. Traditionally this required loading heavy JavaScript/WASM libraries (ZXing, jsQR, Tesseract.js) that add hundreds of kilobytes to bundles and run slowly on mobile devices. The Shape Detection API exposes the operating system's native barcode scanner, face detector, and text recognizer through a simple JavaScript interface — leveraging hardware-accelerated computer vision that the OS already provides for its own features (camera app face detection, Apple's Vision framework, Android's ML Kit). This post covers how the Shape Detection API dispatches to platform-native ML pipelines, what's detectable, and how to build robust scanning experiences.
Problem Statements
-
Platform-Native Detection: How does the Shape Detection API delegate to OS-level computer vision frameworks, and what are the capabilities and limitations of each platform?
-
Barcode Scanning Architecture: How do you build a real-time barcode scanner using
BarcodeDetector, handle multiple barcode formats, and manage camera frame processing? -
Face and Text Detection: How do
FaceDetectorandTextDetectorwork, what features do they expose, and how do they compare to JavaScript alternatives?
Deep Dive: Internal Mechanisms
1. Shape Detection API Architecture
// The API surface:
const barcodeDetector = new BarcodeDetector({ formats: ["qr_code", "ean_13"] });
const faceDetector = new FaceDetector({ maxDetectedFaces: 5, fastMode: true });
const textDetector = new TextDetector();
// All three follow the same pattern:
// 1. Construct a detector
// 2. Pass an ImageBitmapSource (video, canvas, image, blob)
// 3. Get back an array of detected shapes with bounding boxes
/*
* Internal dispatch to platform frameworks:
*
* ┌──────────────────┐ ┌─────────────────────────┐
* │ Shape Detection │ │ Platform Framework │
* │ API (JavaScript) │ ──→ │ │
* │ │ │ macOS/iOS: Vision.framework│
* │ BarcodeDetector │ │ → VNDetectBarcodesRequest│
* │ FaceDetector │ │ → VNDetectFaceRectangles│
* │ TextDetector │ │ → VNRecognizeTextRequest │
* │ │ │ │
* │ │ │ Android: Google ML Kit │
* │ │ │ → BarcodeScannerClient │
* │ │ │ → FaceDetectorClient │
* │ │ │ → TextRecognizerClient │
* │ │ │ │
* │ │ │ Windows: Windows.Media │
* │ │ │ → BarcodeScanner WinRT │
* │ │ │ │
* │ │ │ ChromeOS: ML Service │
* └──────────────────┘ └─────────────────────────┘
*
* Why platform-native matters:
* - Hardware acceleration (GPU, NPU, DSP)
* - Models pre-loaded in OS memory (no download)
* - Optimized for the device's camera characteristics
* - Battery-efficient (dedicated ML silicon)
*/
2. BarcodeDetector: Format Support and Detection
// Check supported formats (varies by platform):
const supportedFormats = await BarcodeDetector.getSupportedFormats();
console.log(supportedFormats);
// Chrome/Android: ["aztec", "codabar", "code_128", "code_39", "code_93",
// "data_matrix", "ean_13", "ean_8", "itf", "pdf417", "qr_code", "upc_a", "upc_e"]
// Chrome/macOS: ["aztec", "code_128", "code_39", "code_93",
// "data_matrix", "ean_13", "ean_8", "itf", "pdf417", "qr_code", "upc_a", "upc_e"]
// Create detector for specific formats:
const detector = new BarcodeDetector({
formats: ["qr_code", "ean_13", "ean_8", "upc_a"],
});
// Detect barcodes in an image:
const image = document.querySelector("img");
const barcodes = await detector.detect(image);
for (const barcode of barcodes) {
console.log(barcode.rawValue); // "https://example.com" (decoded data)
console.log(barcode.format); // "qr_code"
console.log(barcode.boundingBox); // DOMRectReadOnly { x, y, width, height }
console.log(barcode.cornerPoints); // [{x, y}, {x, y}, {x, y}, {x, y}]
/*
* barcode object structure:
* {
* rawValue: string, // Decoded payload
* format: string, // "qr_code", "ean_13", etc.
* boundingBox: DOMRectReadOnly, // Axis-aligned bounding box
* cornerPoints: Point2D[], // Actual corners (may be rotated)
* }
*
* cornerPoints vs boundingBox:
* - boundingBox: always axis-aligned rectangle
* - cornerPoints: actual polygon (handles rotation/perspective)
*
* For a rotated QR code:
* cornerPoints: [{10,50}, {60,10}, {100,60}, {50,100}]
* boundingBox: {x:10, y:10, width:90, height:90} (enclosing rect)
*/
}
3. Real-Time Camera Barcode Scanner
async function createBarcodeScanner(videoElement, onDetected) {
// Feature detection:
if (!("BarcodeDetector" in window)) {
throw new Error("BarcodeDetector not supported");
}
const detector = new BarcodeDetector({
formats: ["qr_code", "ean_13", "ean_8", "upc_a", "upc_e"],
});
// Start camera:
const stream = await navigator.mediaDevices.getUserMedia({
video: {
facingMode: "environment", // Rear camera for scanning
width: { ideal: 1280 },
height: { ideal: 720 },
},
});
videoElement.srcObject = stream;
await videoElement.play();
let scanning = true;
const seenCodes = new Set(); // Deduplicate repeated scans
async function scanFrame() {
if (!scanning) return;
try {
const barcodes = await detector.detect(videoElement);
for (const barcode of barcodes) {
const key = `${barcode.format}:${barcode.rawValue}`;
if (!seenCodes.has(key)) {
seenCodes.add(key);
onDetected(barcode);
}
}
} catch (error) {
// Detection can fail if video frame isn't ready
if (error.name !== "InvalidStateError") {
console.error("Detection error:", error);
}
}
// Throttle to ~15 FPS for detection (don't need 60fps):
// requestAnimationFrame would be excessive
setTimeout(() => requestAnimationFrame(scanFrame), 66);
}
requestAnimationFrame(scanFrame);
return {
stop() {
scanning = false;
stream.getTracks().forEach(t => t.stop());
},
clearHistory() {
seenCodes.clear();
},
};
}
/*
* Scanning performance:
*
* detect() processes one frame at a time.
* Platform-native detection is fast enough for real-time:
*
* ┌──────────────────┬────────────┬───────────────┐
* │ Platform │ Per-frame │ QR decode time │
* ├──────────────────┼────────────┼───────────────┤
* │ Android (ML Kit) │ ~15-30ms │ < 10ms │
* │ macOS (Vision) │ ~20-40ms │ < 15ms │
* │ jsQR (JS lib) │ ~50-150ms │ ~100ms │
* │ ZXing WASM │ ~30-80ms │ ~50ms │
* └──────────────────┴────────────┴───────────────┘
*/
4. FaceDetector: Landmarks and Bounding Boxes
const faceDetector = new FaceDetector({
maxDetectedFaces: 10,
fastMode: false, // true = speed over accuracy
});
const faces = await faceDetector.detect(imageElement);
for (const face of faces) {
console.log(face.boundingBox); // DOMRectReadOnly
// Landmarks (eyes, mouth, nose):
for (const landmark of face.landmarks) {
console.log(landmark.type); // "eye", "mouth", "nose"
console.log(landmark.locations); // [{x, y}] — point coordinates
}
/*
* face object:
* {
* boundingBox: DOMRectReadOnly { x, y, width, height },
* landmarks: [
* { type: "eye", locations: [{x: 100, y: 80}] },
* { type: "eye", locations: [{x: 160, y: 82}] },
* { type: "mouth", locations: [{x: 130, y: 140}] },
* { type: "nose", locations: [{x: 128, y: 110}] },
* ]
* }
*
* Important limitations:
* - NOT face recognition (no identity matching)
* - Detection only: "there is a face here"
* - Landmarks are basic (eyes, mouth, nose centroids)
* - No pose estimation, emotion, or age
* - Landmark availability varies by platform
*
* Platform differences:
* - macOS Vision: good landmark detection
* - Android ML Kit: detailed landmarks when available
* - Some platforms: no landmarks, only boundingBox
*/
}
// Draw face detection overlay:
function drawFaces(ctx, faces, scale) {
ctx.strokeStyle = "#00ff00";
ctx.lineWidth = 2;
for (const face of faces) {
const { x, y, width, height } = face.boundingBox;
ctx.strokeRect(x * scale, y * scale, width * scale, height * scale);
// Draw landmarks:
ctx.fillStyle = "#ff0000";
for (const landmark of face.landmarks) {
for (const point of landmark.locations) {
ctx.beginPath();
ctx.arc(point.x * scale, point.y * scale, 3, 0, Math.PI * 2);
ctx.fill();
}
}
}
}
5. TextDetector: Basic Text Region Detection
// TextDetector finds text regions (not OCR — it doesn't read the text):
const textDetector = new TextDetector();
const textRegions = await textDetector.detect(imageElement);
for (const region of textRegions) {
console.log(region.rawValue); // May be empty (platform-dependent)
console.log(region.boundingBox); // Where text was found
console.log(region.cornerPoints); // Polygon around text
/*
* TextDetector behavior varies significantly by platform:
*
* macOS (Vision): rawValue contains recognized text (actual OCR)
* Android: may only return bounding boxes without rawValue
* Some platforms: TextDetector not supported at all
*
* TextDetector is the least consistently supported detector.
* For reliable OCR, use Tesseract.js or cloud OCR APIs.
*
* TextDetector is most useful for:
* - Finding text regions to crop and send to a server OCR API
* - Highlighting text areas in an AR overlay
* - Guiding user camera positioning ("text detected, hold steady")
*/
}
// Practical pattern: Use TextDetector for ROI, server for OCR:
async function captureAndOCR(videoElement) {
const textDetector = new TextDetector();
const regions = await textDetector.detect(videoElement);
if (regions.length === 0) return null;
// Crop the detected text region:
const region = regions[0];
const canvas = document.createElement("canvas");
const { x, y, width, height } = region.boundingBox;
canvas.width = width;
canvas.height = height;
const ctx = canvas.getContext("2d");
ctx.drawImage(videoElement, x, y, width, height, 0, 0, width, height);
// Send cropped region to server for accurate OCR:
const blob = await new Promise(r => canvas.toBlob(r, "image/png"));
const formData = new FormData();
formData.append("image", blob);
const response = await fetch("/api/ocr", { method: "POST", body: formData });
return response.json();
}
6. Using ImageBitmapSource Inputs Efficiently
// detect() accepts any ImageBitmapSource:
// HTMLImageElement, SVGImageElement, HTMLVideoElement,
// HTMLCanvasElement, ImageBitmap, OffscreenCanvas, VideoFrame, Blob
// Most efficient for video: use VideoFrame (when available):
async function detectFromVideoFrame(detector, videoElement) {
const frame = new VideoFrame(videoElement);
try {
return await detector.detect(frame);
} finally {
frame.close(); // Must close VideoFrame to free memory
}
}
// For batch processing images:
async function detectBarcodesInFiles(files) {
const detector = new BarcodeDetector({ formats: ["qr_code"] });
const results = [];
for (const file of files) {
// Create ImageBitmap (efficient, off-main-thread decoding):
const bitmap = await createImageBitmap(file);
try {
const barcodes = await detector.detect(bitmap);
results.push({ file: file.name, barcodes });
} finally {
bitmap.close(); // Free the bitmap memory
}
}
return results;
}
// Using OffscreenCanvas for preprocessing:
async function scanWithPreprocessing(videoElement) {
const detector = new BarcodeDetector({ formats: ["qr_code"] });
const canvas = new OffscreenCanvas(640, 480);
const ctx = canvas.getContext("2d");
// Draw and enhance:
ctx.drawImage(videoElement, 0, 0, 640, 480);
// Increase contrast for better detection:
ctx.filter = "contrast(1.5) brightness(1.1)";
ctx.drawImage(canvas, 0, 0);
return detector.detect(canvas);
}
7. Barcode Format Deep Dive
/*
* Barcode formats supported by BarcodeDetector:
*
* 1D Barcodes (linear):
* ┌──────────┬──────────────────────────────────┬────────────────┐
* │ Format │ Use Case │ Data │
* ├──────────┼──────────────────────────────────┼────────────────┤
* │ ean_13 │ Products (international) │ 13 digits │
* │ ean_8 │ Small products │ 8 digits │
* │ upc_a │ Products (North America) │ 12 digits │
* │ upc_e │ Small packages │ 8 digits │
* │ code_128 │ Logistics, shipping │ Alphanumeric │
* │ code_39 │ Military, automotive │ A-Z, 0-9, etc. │
* │ code_93 │ Postal, logistics │ Alphanumeric │
* │ codabar │ Libraries, blood banks │ 0-9, -$:/.+ │
* │ itf │ Shipping cartons │ Numeric pairs │
* └──────────┴──────────────────────────────────┴────────────────┘
*
* 2D Barcodes (matrix):
* ┌────────────┬────────────────────────────────┬────────────────┐
* │ Format │ Use Case │ Capacity │
* ├────────────┼────────────────────────────────┼────────────────┤
* │ qr_code │ URLs, contact info, payments │ ~4296 alphanums│
* │ data_matrix│ Manufacturing, electronics │ ~2335 alphanums│
* │ pdf417 │ ID cards, boarding passes │ ~1850 alphanums│
* │ aztec │ Boarding passes (airline) │ ~3832 digits │
* └────────────┴────────────────────────────────┴────────────────┘
*/
// Detecting specific formats for a retail scanner:
const retailScanner = new BarcodeDetector({
formats: ["ean_13", "ean_8", "upc_a", "upc_e"],
});
// Detecting QR codes for a payment/URL scanner:
const qrScanner = new BarcodeDetector({
formats: ["qr_code"],
});
// Tip: Specifying fewer formats is faster than detecting all:
// The detector only runs recognition for specified formats.
8. Handling Multiple Simultaneous Detections
// A single image may contain multiple barcodes:
async function processShelfImage(imageElement) {
const detector = new BarcodeDetector({
formats: ["ean_13", "upc_a"],
});
const barcodes = await detector.detect(imageElement);
// Sort by position (top-to-bottom, left-to-right):
barcodes.sort((a, b) => {
const rowDiff = a.boundingBox.y - b.boundingBox.y;
if (Math.abs(rowDiff) > 50) return rowDiff; // Different row
return a.boundingBox.x - b.boundingBox.x; // Same row, sort by x
});
// Deduplicate (same barcode may be detected slightly shifted):
const unique = [];
for (const barcode of barcodes) {
const isDuplicate = unique.some(existing =>
existing.rawValue === barcode.rawValue &&
Math.abs(existing.boundingBox.x - barcode.boundingBox.x) < 20 &&
Math.abs(existing.boundingBox.y - barcode.boundingBox.y) < 20
);
if (!isDuplicate) {
unique.push(barcode);
}
}
return unique;
}
// Using Intersection Observer pattern for multi-barcode tracking:
function trackBarcodes(videoElement, onEnter, onLeave) {
const detector = new BarcodeDetector({ formats: ["qr_code"] });
const activeCodes = new Map(); // rawValue → last-seen timestamp
const TIMEOUT = 1000; // Consider "left" after 1s not seen
async function scan() {
const barcodes = await detector.detect(videoElement);
const now = Date.now();
// Update seen codes:
for (const barcode of barcodes) {
if (!activeCodes.has(barcode.rawValue)) {
onEnter(barcode); // New barcode entered view
}
activeCodes.set(barcode.rawValue, now);
}
// Check for codes that left the view:
for (const [value, lastSeen] of activeCodes) {
if (now - lastSeen > TIMEOUT) {
activeCodes.delete(value);
onLeave(value);
}
}
requestAnimationFrame(scan);
}
scan();
}
9. Feature Detection and Polyfill Strategy
// Robust feature detection:
async function detectCapabilities() {
const capabilities = {
barcodeDetector: false,
faceDetector: false,
textDetector: false,
supportedBarcodeFormats: [],
};
if ("BarcodeDetector" in window) {
capabilities.barcodeDetector = true;
try {
capabilities.supportedBarcodeFormats =
await BarcodeDetector.getSupportedFormats();
} catch (e) {
capabilities.barcodeDetector = false;
}
}
if ("FaceDetector" in window) {
try {
new FaceDetector();
capabilities.faceDetector = true;
} catch (e) {
// Constructor may throw if platform doesn't support it
}
}
if ("TextDetector" in window) {
try {
new TextDetector();
capabilities.textDetector = true;
} catch (e) {}
}
return capabilities;
}
// Progressive enhancement with fallback:
async function createScanner() {
const caps = await detectCapabilities();
if (caps.barcodeDetector) {
return new NativeBarcodeScanner(); // Uses BarcodeDetector
}
// Fall back to JS library:
const { default: ZXing } = await import("@nicolo-ribaudo/zxing-browser");
return new ZXingScanner(ZXing);
}
10. Integration with Camera Controls and AR Overlays
// Complete barcode scanner with AR overlay:
class ARBarcodeScanner {
constructor(videoElement, overlayCanvas) {
this.video = videoElement;
this.canvas = overlayCanvas;
this.ctx = overlayCanvas.getContext("2d");
this.detector = new BarcodeDetector({
formats: ["qr_code", "ean_13", "ean_8", "upc_a"],
});
this.active = false;
}
async start() {
const stream = await navigator.mediaDevices.getUserMedia({
video: {
facingMode: "environment",
width: { ideal: 1920 },
height: { ideal: 1080 },
},
});
this.video.srcObject = stream;
await this.video.play();
// Match canvas to video dimensions:
this.canvas.width = this.video.videoWidth;
this.canvas.height = this.video.videoHeight;
// Try to enable torch (flashlight) for low-light:
const track = stream.getVideoTracks()[0];
const capabilities = track.getCapabilities();
if (capabilities.torch) {
await track.applyConstraints({ advanced: [{ torch: true }] });
}
// Try to set focus mode to continuous:
if (capabilities.focusMode?.includes("continuous")) {
await track.applyConstraints({
advanced: [{ focusMode: "continuous" }],
});
}
this.active = true;
this.scanLoop();
}
async scanLoop() {
if (!this.active) return;
try {
const barcodes = await this.detector.detect(this.video);
this.drawOverlay(barcodes);
} catch (e) { /* frame not ready */ }
requestAnimationFrame(() => this.scanLoop());
}
drawOverlay(barcodes) {
this.ctx.clearRect(0, 0, this.canvas.width, this.canvas.height);
for (const barcode of barcodes) {
// Draw polygon using cornerPoints:
const points = barcode.cornerPoints;
this.ctx.beginPath();
this.ctx.moveTo(points[0].x, points[0].y);
for (let i = 1; i < points.length; i++) {
this.ctx.lineTo(points[i].x, points[i].y);
}
this.ctx.closePath();
// Highlight:
this.ctx.strokeStyle = "#00ff00";
this.ctx.lineWidth = 3;
this.ctx.stroke();
this.ctx.fillStyle = "rgba(0, 255, 0, 0.1)";
this.ctx.fill();
// Label:
this.ctx.fillStyle = "#00ff00";
this.ctx.font = "16px monospace";
this.ctx.fillText(
`${barcode.format}: ${barcode.rawValue}`,
points[0].x,
points[0].y - 10
);
}
}
stop() {
this.active = false;
const stream = this.video.srcObject;
if (stream) stream.getTracks().forEach(t => t.stop());
}
}
Trade-offs & Considerations
| Approach | Bundle Size | Speed | Formats | Platform Dep. |
|---|---|---|---|---|
| BarcodeDetector (native) | 0 KB | Fastest (hardware-accel) | OS-dependent | High (Chrome only) |
| ZXing (WASM) | ~200 KB | Fast | All standard formats | None |
| jsQR (JS) | ~50 KB | Medium | QR only | None |
| zbar.wasm | ~300 KB | Fast | Most 1D + QR | None |
| FaceDetector (native) | 0 KB | Fastest | N/A (faces) | High |
| face-api.js (TF.js) | ~6 MB models | Slow | N/A (faces + landmarks) | None |
Best Practices
-
Specify only the barcode formats you need in the constructor — the detector runs recognition algorithms for each specified format; restricting to
["qr_code"]is significantly faster than detecting all formats. -
Throttle detection to 10-15 FPS, not every animation frame —
detect()typically takes 15-40ms per call; running at 60fps wastes CPU and battery; 15fps is sufficient for smooth scanning UX. -
Use
cornerPointsrather thanboundingBoxfor AR overlays — cornerPoints provides the actual quadrilateral polygon of the detected shape, correctly handling rotated and perspective-distorted barcodes. -
Always feature-detect with
getSupportedFormats()before constructing —BarcodeDetectormay exist in the global scope but specific formats may not be available on the current platform; the static method is the authoritative capability check. -
Fall back to a JavaScript/WASM library for unsupported platforms — the Shape Detection API is not available in Firefox or older Safari; dynamically import ZXing or jsQR when native detection is unavailable to maintain cross-browser support.
Conclusion
The Shape Detection API exposes platform-native computer vision through three detectors: BarcodeDetector (leveraging OS barcode libraries for 13+ formats including QR, EAN, UPC, Data Matrix), FaceDetector (face bounding boxes and basic landmarks via Vision.framework or ML Kit), and TextDetector (text region detection with inconsistent OCR support). Each detector's detect() method accepts any ImageBitmapSource and returns an array of detected shapes with bounding boxes and corner points. Internally, the browser dispatches to the operating system's ML framework — Apple Vision, Google ML Kit, or Windows Media APIs — which run hardware-accelerated inference on GPU or NPU hardware. This architecture provides zero bundle cost, 2-10x faster detection than JavaScript libraries, and battery-efficient processing. The main trade-off is platform dependency: format support varies by OS, and the API is primarily available in Chromium browsers. The recommended pattern is progressive enhancement: use native detection when available (checking getSupportedFormats() for precise capabilities), with dynamic import of a WASM fallback (ZXing) for unsupported platforms.
What did you think?