Pyyan / Compare / RTMDet vs YOLOv12 vs Grounding DINO

RTMDet vs YOLOv12 vs Grounding DINO

3 of 5

Object Detection · verified 13 Aug 2026

×RTMDetOpenMMLabcurrent
×YOLOv12Ultralyticscurrent
×IGrounding DINOIDEA Researchcurrent
2 slots left
SpecificationRTMDetYOLOv12Grounding DINO
SummaryPure throughput, and an MIT licence.Attention-centric YOLO, still widely deployed.Detects whatever you describe in words.
COCO mAP~52~55~52 zero-shot
Speed300+Very highModerate
LicenceMITAGPL-3.0Apache 2.0
FamilySingle stageSingle stageOpen vocabulary
Open vocabularyNoNoYes
Released202220252023
CategoryObject DetectionObject DetectionObject Detection
OfficialOpenMMLabUltralyticsIDEA Research

Highlighted rows are where these differ.

RTMDet

  • Wins on raw speed where accuracy is sufficient
  • MIT, so no derivative-work obligation

Best for high frame-rate video.

Full spec sheet →

YOLOv12

  • Attention added to the classic single-stage design
  • Same licensing consideration as YOLO26

Best for teams already on the YOLO toolchain.

Full spec sheet →

Grounding DINO

  • Text prompt in, boxes out
  • Fine-tuning still beats it on specialised objects

Best for classes you have no data for.

Full spec sheet →