Florence-2 Model¶
v2¶
Class: Florence2BlockV2 (there are multiple versions of this block)
Source: inference.core.workflows.core_steps.models.foundation.florence2.v2.Florence2BlockV2
Warning: This block has multiple versions. Please refer to the specific version for details. You can learn more about how versions work here: Versioning
Dedicated inference server required (GPU recommended) - you may want to use dedicated deployment
This Workflow block introduces Florence 2, a Visual Language Model (VLM) capable of performing a wide range of tasks, including:
-
Object Detection
-
Instance Segmentation
-
Image Captioning
-
Optical Character Recognition (OCR)
-
and more...
Below is a comprehensive list of tasks supported by the model, along with descriptions on how to utilize their outputs within the Workflows ecosystem:
Task Descriptions:
-
Custom Prompt (
custom) - Use free-form prompt to generate a response. Useful with finetuned models. -
Text Recognition (OCR) (
ocr) - Model recognizes text in the image -
Text Detection & Recognition (OCR) (
ocr-with-text-detection) - Model detects text regions in the image, and then performs OCR on each detected region -
Captioning (short) (
caption) - Model provides a short description of the image -
Captioning (
detailed-caption) - Model provides a long description of the image -
Captioning (long) (
more-detailed-caption) - Model provides a very long description of the image -
Unprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the image -
Object Detection (
open-vocabulary-object-detection) - Model detects and returns the bounding boxes for the provided classes -
Detection & Captioning (
object-detection-and-caption) - Model detects prominent objects and captions them -
Prompted Object Detection (
phrase-grounded-object-detection) - Based on the textual prompt, model detects objects matching the descriptions -
Prompted Instance Segmentation (
phrase-grounded-instance-segmentation) - Based on the textual prompt, model segments objects matching the descriptions -
Segment Bounding Box (
detection-grounded-instance-segmentation) - Model segments the object in the provided bounding box into a polygon -
Classification of Bounding Box (
detection-grounded-classification) - Model classifies the object inside the provided bounding box -
Captioning of Bounding Box (
detection-grounded-caption) - Model captions the object in the provided bounding box -
Text Recognition (OCR) for Bounding Box (
detection-grounded-ocr) - Model performs OCR on the text inside the provided bounding box -
Regions of Interest proposal (
region-proposal) - Model proposes Regions of Interest (Bounding Boxes) in the image
Type identifier¶
Use the following identifier in step "type" field: roboflow_core/florence_2@v2to add the block as
as step in your workflow.
Properties¶
| Name | Type | Description | Refs |
|---|---|---|---|
name |
str |
Enter a unique identifier for this step.. | ❌ |
task_type |
str |
Task type to be performed by model. Value determines required parameters and output response.. | ❌ |
prompt |
str |
Text prompt to the Florence-2 model. | ✅ |
classes |
List[str] |
List of classes to be used. | ✅ |
grounding_detection |
Optional[List[float], List[int]] |
Detection to ground Florence-2 model. May be statically provided bounding box [left_top_x, left_top_y, right_bottom_x, right_bottom_y] or result of object-detection model. If the latter is true, one box will be selected based on grounding_selection_mode.. |
✅ |
grounding_selection_mode |
str |
. | ❌ |
model_id |
str |
Model to be used. | ✅ |
The Refs column marks possibility to parametrise the property with dynamic values available
in workflow runtime. See Bindings for more info.
Runtime compatibility¶
-
hard— runtimeself_hosted_cpu; executionlocal - Requires a GPU; run_locally() loads a model that needs CUDA.
Available Connections¶
Compatible Blocks
Check what blocks you can connect to Florence-2 Model in version v2.
- inputs:
SAM 3,Image Preprocessing,Anthropic Claude,Image Slicer,Time in Zone,Dynamic Crop,Mask Area Measurement,BoT-SORT Tracker,Bounding Box Visualization,Mask Edge Snap,Object Detection Model,Path Deviation,Absolute Static Crop,SIFT Comparison,Stitch Images,Stitch OCR Detections,OpenAI,Instance Segmentation Model,Email Notification,Stability AI Inpainting,EasyOCR,Llama 3.2 Vision,Track Class Lock,Florence-2 Model,Gaze Detection,Roboflow Custom Metadata,Dynamic Zone,Auto Rotate on Edges,YOLO-World Model,Keypoint Detection Model,Detections Transformation,Byte Tracker,Stability AI Outpainting,Model Comparison Visualization,Slack Notification,Detections Classes Replacement,Line Counter Visualization,Camera Calibration,Byte Tracker,Single-Label Classification Model,Clip Comparison,VLM As Detector,CogVLM,SORT Tracker,Camera Focus,Corner Visualization,Ellipse Visualization,PP-OCR,Morphological Transformation,Detections List Roll-Up,Anthropic Claude,Roboflow Visual Search,Color Visualization,Instance Segmentation Model,OpenAI,Triangle Visualization,Time in Zone,Detections Stabilizer,Detection Event Log,Object Detection Model,Image Contours,SAM 3,Image Threshold,SAM 3,Detections Merge,Current Time,Roboflow Visual Search Classifier,QR Code Generator,OpenAI-Compatible LLM,Florence-2 Model,Polygon Zone Visualization,Stitch OCR Detections,Roboflow Asset Library Attributes,GeoTag Detection,Path Deviation,Microsoft SQL Server Sink,Moondream2,VLM As Classifier,Camera Focus,Image Convert Grayscale,Label Visualization,Llama 3.2 Vision,Stability AI Image Generation,Detection Offset,VLM As Detector,Instance Segmentation Model,Line Counter,MQTT Writer,Roboflow Dataset Upload,Local File Sink,Bounding Rectangle,Google Gemini,Event Writer,Depth Estimation,Seg Preview,Byte Tracker,Google Gemini,OpenAI,Trace Visualization,Twilio SMS Notification,SAM 3 Interactive,PLC EthernetIP,Detections Filter,LMM For Classification,Object Detection Model,Velocity,Webhook Sink,Halo Visualization,Buffer,Mask Visualization,Template Matching,Pixelate Visualization,Twilio SMS/MMS Notification,MoonshotAI Kimi,Dot Visualization,Multi-Label Classification Model,Image Stack,OPC UA Writer Sink,Google Gemini,OC-SORT Tracker,Keypoint Visualization,Dimension Collapse,LMM,Detections Combine,Image Slicer,PTZ Tracking (ONVIF),Time in Zone,OCR Model,Circle Visualization,Contrast Enhancement,Relative Static Crop,SAM2 Video Tracker,ByteTrack Tracker,Morphological Transformation,Per-Class Confidence Filter,Email Notification,Halo Visualization,Clip Comparison,Cosmos 3,Polygon Visualization,Qwen-VL,PLC Writer,Google Gemma,Crop Visualization,Qwen 3.5 API,Model Monitoring Inference Aggregator,Keypoint Detection Model,Size Measurement,Icon Visualization,Heatmap Visualization,Motion Detection,Google Gemma API,Detections Consensus,Instance Segmentation Model,CSV Formatter,Image Blur,Segment Anything 2 Model,Background Color Visualization,Grid Visualization,SAM3 Video Tracker,Detections Stitch,Blur Visualization,GLM-OCR,Anthropic Claude,Reference Path Visualization,Classification Label Visualization,Google Vision OCR,Perspective Correction,Background Subtraction,Polygon Visualization,Contrast Equalization,SIFT,Qwen 3.6 API,Text Display,MoonshotAI Kimi,OpenRouter,Qwen3.5-VL,Roboflow Vision Events,OpenAI,Keypoint Detection Model,PLC ModbusTCP,Overlap Filter,S3 Sink,Roboflow Dataset Upload - outputs:
SAM 3,Image Preprocessing,Anthropic Claude,Time in Zone,CLIP Embedding Model,Dynamic Crop,Bounding Box Visualization,Cache Get,Object Detection Model,Path Deviation,SIFT Comparison,Stitch OCR Detections,OpenAI,Instance Segmentation Model,Email Notification,Stability AI Inpainting,Distance Measurement,Frame Delay,Llama 3.2 Vision,Florence-2 Model,Roboflow Custom Metadata,YOLO-World Model,Auto Rotate on Edges,Keypoint Detection Model,Stability AI Outpainting,Model Comparison Visualization,Slack Notification,Detections Classes Replacement,Line Counter Visualization,PLC Reader,Clip Comparison,VLM As Detector,CogVLM,Corner Visualization,Ellipse Visualization,Morphological Transformation,Detections List Roll-Up,Anthropic Claude,Roboflow Visual Search,Color Visualization,Instance Segmentation Model,OpenAI,Triangle Visualization,Time in Zone,Object Detection Model,SAM 3,Image Threshold,SAM 3,Current Time,Roboflow Visual Search Classifier,QR Code Generator,OpenAI-Compatible LLM,Florence-2 Model,Semantic Segmentation Model,Polygon Zone Visualization,Stitch OCR Detections,Path Deviation,Roboflow Asset Library Attributes,Microsoft SQL Server Sink,Moondream2,VLM As Classifier,Label Visualization,Llama 3.2 Vision,Stability AI Image Generation,VLM As Detector,Instance Segmentation Model,Perception Encoder Embedding Model,Line Counter,MQTT Writer,Roboflow Dataset Upload,Local File Sink,Cache Set,VLM As Classifier,Event Writer,Google Gemini,Depth Estimation,Seg Preview,Google Gemini,OpenAI,Trace Visualization,Twilio SMS Notification,Pixel Color Count,PLC EthernetIP,LMM For Classification,Object Detection Model,Webhook Sink,Halo Visualization,Buffer,Mask Visualization,Twilio SMS/MMS Notification,MoonshotAI Kimi,Dot Visualization,OPC UA Writer Sink,Google Gemini,Keypoint Visualization,LMM,PTZ Tracking (ONVIF),Time in Zone,Circle Visualization,Morphological Transformation,Per-Class Confidence Filter,Email Notification,Halo Visualization,Clip Comparison,Cosmos 3,Polygon Visualization,Qwen-VL,Google Gemma,Crop Visualization,Qwen 3.5 API,Model Monitoring Inference Aggregator,Keypoint Detection Model,Size Measurement,Icon Visualization,Heatmap Visualization,Single-Label Classification Model,Motion Detection,Google Gemma API,Instance Segmentation Model,Detections Consensus,Segment Anything 2 Model,Image Blur,Background Color Visualization,SAM3 Video Tracker,Detections Stitch,Grid Visualization,GLM-OCR,Anthropic Claude,Reference Path Visualization,Classification Label Visualization,Google Vision OCR,Perspective Correction,Polygon Visualization,Contrast Equalization,Qwen 3.6 API,Text Display,JSON Parser,MoonshotAI Kimi,OpenRouter,Qwen3.5-VL,Roboflow Vision Events,Keypoint Detection Model,OpenAI,Multi-Label Classification Model,S3 Sink,Line Counter,Roboflow Dataset Upload
Input and Output Bindings¶
The available connections depend on its binding kinds. Check what binding kinds
Florence-2 Model in version v2 has.
Bindings
-
input
images(image): The image to infer on..prompt(string): Text prompt to the Florence-2 model.classes(list_of_values): List of classes to be used.grounding_detection(Union[instance_segmentation_prediction,object_detection_prediction,list_of_values,keypoint_detection_prediction]): Detection to ground Florence-2 model. May be statically provided bounding box[left_top_x, left_top_y, right_bottom_x, right_bottom_y]or result of object-detection model. If the latter is true, one box will be selected based ongrounding_selection_mode..model_id(roboflow_model_id): Model to be used.
-
output
raw_output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.parsed_output(dictionary): Dictionary.classes(list_of_values): List of values of any type.
Example JSON definition of step Florence-2 Model in version v2
{
"name": "<your_step_name_here>",
"type": "roboflow_core/florence_2@v2",
"images": "$inputs.image",
"task_type": "<block_does_not_provide_example>",
"prompt": "my prompt",
"classes": [
"class-a",
"class-b"
],
"grounding_detection": "$steps.detection.predictions",
"grounding_selection_mode": "first",
"model_id": "florence-2-base"
}
v1¶
Class: Florence2BlockV1 (there are multiple versions of this block)
Source: inference.core.workflows.core_steps.models.foundation.florence2.v1.Florence2BlockV1
Warning: This block has multiple versions. Please refer to the specific version for details. You can learn more about how versions work here: Versioning
Dedicated inference server required (GPU recommended) - you may want to use dedicated deployment
This Workflow block introduces Florence 2, a Visual Language Model (VLM) capable of performing a wide range of tasks, including:
-
Object Detection
-
Instance Segmentation
-
Image Captioning
-
Optical Character Recognition (OCR)
-
and more...
Below is a comprehensive list of tasks supported by the model, along with descriptions on how to utilize their outputs within the Workflows ecosystem:
Task Descriptions:
-
Custom Prompt (
custom) - Use free-form prompt to generate a response. Useful with finetuned models. -
Text Recognition (OCR) (
ocr) - Model recognizes text in the image -
Text Detection & Recognition (OCR) (
ocr-with-text-detection) - Model detects text regions in the image, and then performs OCR on each detected region -
Captioning (short) (
caption) - Model provides a short description of the image -
Captioning (
detailed-caption) - Model provides a long description of the image -
Captioning (long) (
more-detailed-caption) - Model provides a very long description of the image -
Unprompted Object Detection (
object-detection) - Model detects and returns the bounding boxes for prominent objects in the image -
Object Detection (
open-vocabulary-object-detection) - Model detects and returns the bounding boxes for the provided classes -
Detection & Captioning (
object-detection-and-caption) - Model detects prominent objects and captions them -
Prompted Object Detection (
phrase-grounded-object-detection) - Based on the textual prompt, model detects objects matching the descriptions -
Prompted Instance Segmentation (
phrase-grounded-instance-segmentation) - Based on the textual prompt, model segments objects matching the descriptions -
Segment Bounding Box (
detection-grounded-instance-segmentation) - Model segments the object in the provided bounding box into a polygon -
Classification of Bounding Box (
detection-grounded-classification) - Model classifies the object inside the provided bounding box -
Captioning of Bounding Box (
detection-grounded-caption) - Model captions the object in the provided bounding box -
Text Recognition (OCR) for Bounding Box (
detection-grounded-ocr) - Model performs OCR on the text inside the provided bounding box -
Regions of Interest proposal (
region-proposal) - Model proposes Regions of Interest (Bounding Boxes) in the image
Type identifier¶
Use the following identifier in step "type" field: roboflow_core/florence_2@v1to add the block as
as step in your workflow.
Properties¶
| Name | Type | Description | Refs |
|---|---|---|---|
name |
str |
Enter a unique identifier for this step.. | ❌ |
task_type |
str |
Task type to be performed by model. Value determines required parameters and output response.. | ❌ |
prompt |
str |
Text prompt to the Florence-2 model. | ✅ |
classes |
List[str] |
List of classes to be used. | ✅ |
grounding_detection |
Optional[List[float], List[int]] |
Detection to ground Florence-2 model. May be statically provided bounding box [left_top_x, left_top_y, right_bottom_x, right_bottom_y] or result of object-detection model. If the latter is true, one box will be selected based on grounding_selection_mode.. |
✅ |
grounding_selection_mode |
str |
. | ❌ |
model_version |
str |
Model to be used. | ✅ |
The Refs column marks possibility to parametrise the property with dynamic values available
in workflow runtime. See Bindings for more info.
Runtime compatibility¶
-
hard— runtimeself_hosted_cpu; executionlocal - Requires a GPU; run_locally() loads a model that needs CUDA.
Available Connections¶
Compatible Blocks
Check what blocks you can connect to Florence-2 Model in version v1.
- inputs:
SAM 3,Image Preprocessing,Anthropic Claude,Image Slicer,Time in Zone,Dynamic Crop,Mask Area Measurement,BoT-SORT Tracker,Bounding Box Visualization,Mask Edge Snap,Object Detection Model,Path Deviation,Absolute Static Crop,SIFT Comparison,Stitch Images,Stitch OCR Detections,OpenAI,Instance Segmentation Model,Email Notification,Stability AI Inpainting,EasyOCR,Llama 3.2 Vision,Track Class Lock,Florence-2 Model,Gaze Detection,Roboflow Custom Metadata,Dynamic Zone,Auto Rotate on Edges,YOLO-World Model,Keypoint Detection Model,Detections Transformation,Byte Tracker,Stability AI Outpainting,Model Comparison Visualization,Slack Notification,Detections Classes Replacement,Line Counter Visualization,Camera Calibration,Byte Tracker,Single-Label Classification Model,Clip Comparison,VLM As Detector,CogVLM,SORT Tracker,Camera Focus,Corner Visualization,Ellipse Visualization,PP-OCR,Morphological Transformation,Detections List Roll-Up,Anthropic Claude,Roboflow Visual Search,Color Visualization,Instance Segmentation Model,OpenAI,Triangle Visualization,Time in Zone,Detections Stabilizer,Detection Event Log,Object Detection Model,Image Contours,SAM 3,Image Threshold,SAM 3,Detections Merge,Current Time,Roboflow Visual Search Classifier,QR Code Generator,OpenAI-Compatible LLM,Florence-2 Model,Polygon Zone Visualization,Stitch OCR Detections,Roboflow Asset Library Attributes,GeoTag Detection,Path Deviation,Microsoft SQL Server Sink,Moondream2,VLM As Classifier,Camera Focus,Image Convert Grayscale,Label Visualization,Llama 3.2 Vision,Stability AI Image Generation,Detection Offset,VLM As Detector,Instance Segmentation Model,Line Counter,MQTT Writer,Roboflow Dataset Upload,Local File Sink,Bounding Rectangle,Google Gemini,Event Writer,Depth Estimation,Seg Preview,Byte Tracker,Google Gemini,OpenAI,Trace Visualization,Twilio SMS Notification,SAM 3 Interactive,PLC EthernetIP,Detections Filter,LMM For Classification,Object Detection Model,Velocity,Webhook Sink,Halo Visualization,Buffer,Mask Visualization,Template Matching,Pixelate Visualization,Twilio SMS/MMS Notification,MoonshotAI Kimi,Dot Visualization,Multi-Label Classification Model,Image Stack,OPC UA Writer Sink,Google Gemini,OC-SORT Tracker,Keypoint Visualization,Dimension Collapse,LMM,Detections Combine,Image Slicer,PTZ Tracking (ONVIF),Time in Zone,OCR Model,Circle Visualization,Contrast Enhancement,Relative Static Crop,SAM2 Video Tracker,ByteTrack Tracker,Morphological Transformation,Per-Class Confidence Filter,Email Notification,Halo Visualization,Clip Comparison,Cosmos 3,Polygon Visualization,Qwen-VL,PLC Writer,Google Gemma,Crop Visualization,Qwen 3.5 API,Model Monitoring Inference Aggregator,Keypoint Detection Model,Size Measurement,Icon Visualization,Heatmap Visualization,Motion Detection,Google Gemma API,Detections Consensus,Instance Segmentation Model,CSV Formatter,Image Blur,Segment Anything 2 Model,Background Color Visualization,Grid Visualization,SAM3 Video Tracker,Detections Stitch,Blur Visualization,GLM-OCR,Anthropic Claude,Reference Path Visualization,Classification Label Visualization,Google Vision OCR,Perspective Correction,Background Subtraction,Polygon Visualization,Contrast Equalization,SIFT,Qwen 3.6 API,Text Display,MoonshotAI Kimi,OpenRouter,Qwen3.5-VL,Roboflow Vision Events,OpenAI,Keypoint Detection Model,PLC ModbusTCP,Overlap Filter,S3 Sink,Roboflow Dataset Upload - outputs:
SAM 3,Image Preprocessing,Anthropic Claude,Time in Zone,CLIP Embedding Model,Dynamic Crop,Bounding Box Visualization,Cache Get,Object Detection Model,Path Deviation,SIFT Comparison,Stitch OCR Detections,OpenAI,Instance Segmentation Model,Email Notification,Stability AI Inpainting,Distance Measurement,Frame Delay,Llama 3.2 Vision,Florence-2 Model,Roboflow Custom Metadata,YOLO-World Model,Auto Rotate on Edges,Keypoint Detection Model,Stability AI Outpainting,Model Comparison Visualization,Slack Notification,Detections Classes Replacement,Line Counter Visualization,PLC Reader,Clip Comparison,VLM As Detector,CogVLM,Corner Visualization,Ellipse Visualization,Morphological Transformation,Detections List Roll-Up,Anthropic Claude,Roboflow Visual Search,Color Visualization,Instance Segmentation Model,OpenAI,Triangle Visualization,Time in Zone,Object Detection Model,SAM 3,Image Threshold,SAM 3,Current Time,Roboflow Visual Search Classifier,QR Code Generator,OpenAI-Compatible LLM,Florence-2 Model,Semantic Segmentation Model,Polygon Zone Visualization,Stitch OCR Detections,Path Deviation,Roboflow Asset Library Attributes,Microsoft SQL Server Sink,Moondream2,VLM As Classifier,Label Visualization,Llama 3.2 Vision,Stability AI Image Generation,VLM As Detector,Instance Segmentation Model,Perception Encoder Embedding Model,Line Counter,MQTT Writer,Roboflow Dataset Upload,Local File Sink,Cache Set,VLM As Classifier,Event Writer,Google Gemini,Depth Estimation,Seg Preview,Google Gemini,OpenAI,Trace Visualization,Twilio SMS Notification,Pixel Color Count,PLC EthernetIP,LMM For Classification,Object Detection Model,Webhook Sink,Halo Visualization,Buffer,Mask Visualization,Twilio SMS/MMS Notification,MoonshotAI Kimi,Dot Visualization,OPC UA Writer Sink,Google Gemini,Keypoint Visualization,LMM,PTZ Tracking (ONVIF),Time in Zone,Circle Visualization,Morphological Transformation,Per-Class Confidence Filter,Email Notification,Halo Visualization,Clip Comparison,Cosmos 3,Polygon Visualization,Qwen-VL,Google Gemma,Crop Visualization,Qwen 3.5 API,Model Monitoring Inference Aggregator,Keypoint Detection Model,Size Measurement,Icon Visualization,Heatmap Visualization,Single-Label Classification Model,Motion Detection,Google Gemma API,Instance Segmentation Model,Detections Consensus,Segment Anything 2 Model,Image Blur,Background Color Visualization,SAM3 Video Tracker,Detections Stitch,Grid Visualization,GLM-OCR,Anthropic Claude,Reference Path Visualization,Classification Label Visualization,Google Vision OCR,Perspective Correction,Polygon Visualization,Contrast Equalization,Qwen 3.6 API,Text Display,JSON Parser,MoonshotAI Kimi,OpenRouter,Qwen3.5-VL,Roboflow Vision Events,Keypoint Detection Model,OpenAI,Multi-Label Classification Model,S3 Sink,Line Counter,Roboflow Dataset Upload
Input and Output Bindings¶
The available connections depend on its binding kinds. Check what binding kinds
Florence-2 Model in version v1 has.
Bindings
-
input
images(image): The image to infer on..prompt(string): Text prompt to the Florence-2 model.classes(list_of_values): List of classes to be used.grounding_detection(Union[instance_segmentation_prediction,object_detection_prediction,list_of_values,keypoint_detection_prediction]): Detection to ground Florence-2 model. May be statically provided bounding box[left_top_x, left_top_y, right_bottom_x, right_bottom_y]or result of object-detection model. If the latter is true, one box will be selected based ongrounding_selection_mode..model_version(string): Model to be used.
-
output
raw_output(Union[string,language_model_output]): String value ifstringor LLM / VLM output iflanguage_model_output.parsed_output(dictionary): Dictionary.classes(list_of_values): List of values of any type.
Example JSON definition of step Florence-2 Model in version v1
{
"name": "<your_step_name_here>",
"type": "roboflow_core/florence_2@v1",
"images": "$inputs.image",
"task_type": "<block_does_not_provide_example>",
"prompt": "my prompt",
"classes": [
"class-a",
"class-b"
],
"grounding_detection": "$steps.detection.predictions",
"grounding_selection_mode": "first",
"model_version": "florence-2-base"
}