Vision
Computer vision models and pipelines — recognition, detection and image understanding beyond documents.

Turns any PDF into a fillable form: FFDNet models detect text, checkbox and signature fields; one CLI command writes the interactive PDF. Paper, dataset and weights all open.
Lightweight Python face recognition and facial-attribute analysis (age, gender, emotion) wrapping VGG-Face, FaceNet, ArcFace and friends — pip install, self-hosted, battle-tested at 23k stars.
Python package bridging deep learning and geospatial data: train and apply classification, detection and segmentation models on satellite and aerial imagery. JOSS paper, conda-forge, QGIS plugin.
Berkeley's visual RAG: render pages and PDFs to screenshot tiles and retrieve with a VLM embedder — tables, charts and layout survive. pixelshot CLI plus a hosted 8.28M-page Wikipedia index.

Meta's Segment Anything 2: promptable zero-shot segmentation for images and video with streaming memory — click and box prompts become tracked masks in real time.

All-in-one server and WebUI for generative image and video — Stable Diffusion and dozens of model families, with captioning, upscaling and processing pipelines. Cross-platform, API-first.

Roboflow's reusable computer-vision toolkit: one Detections API over any model (YOLO, SAM, transformers), 20+ annotators, zone counting, tracking and dataset tools. 48k stars, MIT.

Ultralytics YOLO (v8→26): real-time object detection, segmentation, classification, pose and tracking behind one Python/CLI API — train, validate and export to ONNX/TensorRT/CoreML.

All-in-one agentic framework for video: understanding and summarization, clip editing, and generative remaking, driven end-to-end through natural-language conversation.
HKUDS multi-agent video studio: turns an idea, novel or screenplay into a finished film — scriptwriting, storyboards, consistent characters, then rendering via Seedance/Nano Banana/Omni APIs.