Image Stack¶
Class: ImageStackBlockV1
Source: inference.core.workflows.core_steps.fusion.image_stack.v1.ImageStackBlockV1
Accumulate compressed video frames into a fixed-size stack, returning the most recent N frames as JPEG-encoded binary blobs. Designed for shared-hosting safety: frames are always JPEG-compressed and downsampled to fit within resolution limits, preventing out-of-memory conditions.
How This Block Works¶
- Receives a video frame (WorkflowImageData) each workflow cycle.
- Downsamples the frame if it exceeds the configured resolution limits (default 1920x1080), preserving aspect ratio.
- JPEG-encodes the frame at quality 75 and stores the resulting bytes.
- Maintains a per-camera FIFO buffer (deque) of up to
stack_sizecompressed frames. When the buffer is full the oldest frame is automatically evicted. - If
stack_sizechanges between calls (e.g. via a dynamic selector), the buffer is resized and existing frames are preserved up to the new limit. - If the
clearinput is True the buffer is flushed before the current frame is added. - Outputs the list of JPEG byte blobs (newest first) and the current frame count.
Common Use Cases¶
- Action / activity recognition: accumulate a clip of N frames and pass them to a vision-language model (e.g. Google Gemini, Qwen) that can reason over multiple images to classify actions, detect events, or describe what is happening in a scene.
- Time-lapse snapshots: collect the last N frames for periodic visual comparison.
- Event buffering: keep a rolling window of frames around an event of interest.
Type identifier¶
Use the following identifier in step "type" field: roboflow_core/image_stack@v1to add the block as
as step in your workflow.
Properties¶
| Name | Type | Description | Refs |
|---|---|---|---|
name |
str |
Enter a unique identifier for this step.. | ❌ |
stack_size |
int |
Maximum number of frames to keep in the stack (1-64). When the stack is full the oldest frame is evicted.. | ✅ |
resolution_width |
int |
Maximum frame width in pixels (64-1920). Frames wider than this are downsampled preserving aspect ratio.. | ✅ |
resolution_height |
int |
Maximum frame height in pixels (64-1080). Frames taller than this are downsampled preserving aspect ratio.. | ✅ |
clear |
bool |
When True the entire frame buffer is flushed before the current frame is added. Useful for resetting state on scene changes.. | ✅ |
The Refs column marks possibility to parametrise the property with dynamic values available
in workflow runtime. See Bindings for more info.
Runtime compatibility¶
-
soft— runtimehosted_serverless,dedicated_deployment; executionremote; inputvideo - Frame stack is stored in process memory per video_identifier. With remote step execution on stateless or multi-replica HTTP runtimes, successive frames may be served by different worker processes, so the stack resets or contains only a partial frame history. Use local step execution in an InferencePipeline for stable cross-frame results.
-
soft— inputimage - Block depends on temporal context from video or repeated-frame workflows. With a still image/photo, there is no meaningful history to track, compare, aggregate, or visualize, so the block provides little or no benefit.
Available Connections¶
Compatible Blocks
Check what blocks you can connect to Image Stack in version v1.
- inputs:
PLC Writer,Image Blur,Stability AI Outpainting,Blur Visualization,Frame Delay,Bounding Box Visualization,Absolute Static Crop,Crop Visualization,Reference Path Visualization,Roboflow Dataset Upload,Polygon Visualization,VLM As Detector,Roboflow Visual Search Classifier,Event Writer,Image Slicer,Image Convert Grayscale,VLM As Classifier,Email Notification,Webhook Sink,Contrast Enhancement,Template Matching,VLM As Classifier,Perspective Correction,Motion Detection,Halo Visualization,SIFT,Slack Notification,Label Visualization,Dynamic Crop,PLC Reader,Detections Consensus,Label Visualization,S3 Sink,Halo Visualization,Line Counter Visualization,Email Notification,Detection Event Log,Roboflow Asset Library Attributes,Pixel Color Count,Background Subtraction,Ellipse Visualization,Stitch Images,OPC UA Writer Sink,Relative Static Crop,Corner Visualization,Triangle Visualization,Depth Estimation,SIFT Comparison,JSON Parser,Distance Measurement,Dot Visualization,Camera Calibration,Background Color Visualization,Polygon Zone Visualization,Image Contours,Roboflow Custom Metadata,Roboflow Dataset Upload,Camera Focus,Stability AI Inpainting,Contrast Equalization,Morphological Transformation,SIFT Comparison,Image Threshold,Trace Visualization,Color Visualization,Camera Focus,Line Counter,Stability AI Image Generation,Local File Sink,Circle Visualization,QR Code Generator,Morphological Transformation,Dynamic Zone,Model Monitoring Inference Aggregator,Roboflow Vision Events,Heatmap Visualization,Twilio SMS Notification,Auto Rotate on Edges,Polygon Visualization,Microsoft SQL Server Sink,MQTT Writer,Icon Visualization,Rich Label Visualization,Identify Changes,Twilio SMS/MMS Notification,Keypoint Visualization,Image Preprocessing,Image Slicer,Line Counter,Text Display,Image Stack,Grid Visualization,Model Comparison Visualization,PTZ Tracking (ONVIF),Identify Outliers,Mask Visualization,Classification Label Visualization,Pixelate Visualization,Roboflow Visual Search,VLM As Detector - outputs:
Track Class Lock,Image Blur,Byte Tracker,Mask Edge Snap,Path Deviation,Crop Visualization,VLM As Detector,Object Detection Model,Polygon Visualization,Roboflow Visual Search Classifier,Detections Stabilizer,Image Slicer,Webhook Sink,VLM As Classifier,Motion Detection,Keypoint Detection Model,Clip Comparison,YOLO-World Model,Buffer,OC-SORT Tracker,Slack Notification,Label Visualization,MoonshotAI Kimi,PLC Reader,Stitch OCR Detections,Label Visualization,Instance Segmentation Model,Keypoint Detection Model,Email Notification,BoT-SORT Tracker,Pixel Color Count,OPC UA Writer Sink,Ellipse Visualization,Path Deviation,Background Subtraction,Stitch Images,Corner Visualization,Triangle Visualization,Qwen-VL,Detection Offset,Seg Preview,Google Gemini,Polygon Zone Visualization,LMM For Classification,Trace Visualization,SIFT Comparison,Object Detection Model,Color Visualization,Llama 3.2 Vision,SAM3 Video Tracker,Llama 3.2 Vision,Dominant Color,Morphological Transformation,Dynamic Zone,SAM2 Video Tracker,Google Gemma API,Google Gemma,Polygon Visualization,Icon Visualization,Rich Label Visualization,Identify Changes,Twilio SMS/MMS Notification,Keypoint Visualization,OpenAI,Image Slicer,Clip Comparison,Keypoint Detection Model,Object Detection Model,Qwen 3.6 API,Instance Segmentation Model,Mask Visualization,Nearest Neighbor Detection Match,OpenAI,Byte Tracker,Instance Segmentation Model,Roboflow Visual Search,VLM As Detector,Stability AI Outpainting,Blur Visualization,Frame Delay,Bounding Box Visualization,Roboflow Dataset Upload,Absolute Static Crop,Reference Path Visualization,OpenAI,Anthropic Claude,Google Gemini,Event Writer,Cache Set,Stitch OCR Detections,Time in Zone,VLM As Classifier,Email Notification,Perspective Correction,Detections List Roll-Up,Halo Visualization,OpenRouter,Detections Consensus,PLC EthernetIP,Halo Visualization,Line Counter Visualization,Time in Zone,Time in Zone,Roboflow Asset Library Attributes,Size Measurement,SIFT Comparison,SAM 3,ByteTrack Tracker,Dot Visualization,Florence-2 Model,Google Gemini,Image Contours,Roboflow Dataset Upload,MoonshotAI Kimi,Qwen 3.5 API,Stability AI Inpainting,Morphological Transformation,Image Threshold,SAM 3,Line Counter,Instance Segmentation Model,Circle Visualization,QR Code Generator,Google Gemini,Detections Classes Replacement,Roboflow Vision Events,Heatmap Visualization,Twilio SMS Notification,Auto Rotate on Edges,MQTT Writer,SAM 3,Anthropic Claude,SORT Tracker,Image Preprocessing,Line Counter,Anthropic Claude,Text Display,Image Stack,Grid Visualization,Florence-2 Model,PTZ Tracking (ONVIF),Identify Outliers,Classification Label Visualization,Byte Tracker,Pixelate Visualization
Input and Output Bindings¶
The available connections depend on its binding kinds. Check what binding kinds
Image Stack in version v1 has.
Bindings
-
input
image(image): Video frame to add to the stack..stack_size(integer): Maximum number of frames to keep in the stack (1-64). When the stack is full the oldest frame is evicted..resolution_width(integer): Maximum frame width in pixels (64-1920). Frames wider than this are downsampled preserving aspect ratio..resolution_height(integer): Maximum frame height in pixels (64-1080). Frames taller than this are downsampled preserving aspect ratio..clear(boolean): When True the entire frame buffer is flushed before the current frame is added. Useful for resetting state on scene changes..
-
output
frames(list_of_values): List of values of any type.frames_count(integer): Integer value.
Example JSON definition of step Image Stack in version v1
{
"name": "<your_step_name_here>",
"type": "roboflow_core/image_stack@v1",
"image": "$inputs.image",
"stack_size": 5,
"resolution_width": 640,
"resolution_height": 480,
"clear": false
}