This microservice demonstrates AI-powered captioning for multiple live RTSP video streams using Intel® DL Streamer and OpenVINO™ Vision Language Models (VLMs). It enables real-time understanding and natural language description of video content, showcasing how modern vision-language AI can be applied to streaming video analytics pipelines.
🎥 Live RTSP Ingestion – Continuous processing of IP camera or streaming sources
🧠 Vision‑Language Captioning – Prompt‑driven captions via OpenVINO VLMs
⚡ Intel‑Optimized Inference – Low‑latency, high‑throughput execution
📡 WebRTC Preview – Real‑time video preview alongside captions
📊 Performance Metrics – CPU/GPU/RAM usage and inference stats
🔄 Multi‑Model Support – Easy switching between supported VLMs
For more details on deployment, refer to the documentation.
Copyright (C) 2025 Intel Corporation.
Licensed under the Apache License, Version 2.0 (the "License"); you may not use this file except in compliance with the License. You may obtain a copy of the License at http://www.apache.org/licenses/LICENSE-2.0
Intel, the Intel logo, and Xeon are trademarks of Intel Corporation in the U.S. and/or other countries.
*Other names and brands may be claimed as the property of others.
Content type
Image
Digest
sha256:5dc9cd47d…
Size
64.6 MB
Last updated
5 days ago
docker pull intel/video-caption-servicePulls:
155
Sep 14 to Sep 20