AUTO-SYNC Index refreshed every 12h · Evidence-linked
Data release v20260914_020943 Generated 2026-09-14 Methodology Report missing resource
Data checked 2026-08-13. Data may be stale - beyond the review cycle. Review cycles are documented on the Methodology page. Methodology

IDEA-Research/groundingdino

AI Model Tier A perception_ai Apache-2.0
Official confirmed

[ECCV 2024] Official implementation of the paper "Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection"

Engineering Snapshot

Use cases & tasks 4
Best suited for
  • UAV perception requiring zero-shot detection of arbitrary objects
  • Vision-language grounding of text commands to bounding boxes
Primary tasks
  • Text-prompted object detection from drone camera feeds
  • Open-vocabulary scene understanding for low-altitude AI
Stack & ecosystem
Resource type
AI Model
Ecosystem
object-detection · open-world · open-world-detection · vision-language · vision-language-transformer
License & compliance
License
Apache-2.0 (Inferred)
Lifecycle & freshness Show
Maintenance
Active
Latest version
Not recorded
Last activity
2026-08-13
Last checked
2026-08-13
Verification
Official confirmed

What It Solves

Enable flexible open-set object detection in UAV/low-altitude imagery using natural language prompts without retraining for closed-set classes.

Primary use cases

  • Text-prompted object detection from drone camera feeds
  • Open-vocabulary scene understanding for low-altitude AI

Secondary use cases

  • Research reproduction of ECCV 2024 open-set detection method

When to Use

Consider when

  • Target classes are not known a priori
  • Natural language descriptions of objects are available

Verify before adopting

  • No official ONNX/TensorRT export documented
  • Training data composition not fully disclosed
  • Two-stage encoder latency vs closed-set detectors

Start Here

github api https://api.github.com/repos/IDEA-Research/groundingdino

Adoption Checklist

  • Needs verification No official ONNX/TensorRT export documented
  • Needs verification Training data composition not fully disclosed
  • Needs verification Two-stage encoder latency vs closed-set detectors

Each check stays "needs verification" until an official source confirms it; unconfirmed items are never marked verified.

Known Limitations & Unknowns

Known limitations

  • Two-stage encoder (image + text) increases latency vs closed-set detectors
  • Performance depends on text prompt quality
  • No official ONNX/TensorRT export documented in source
  • Training data composition not fully disclosed in repo

Not publicly verified

  • Framework
  • Model size
  • Runtime
  • Edge feasibility evidence
  • Benchmark source
  • Specific training data

Alternatives & Related Tools

Alternatives

How is it used?

Start from the recorded entry points below, then validate against the technical checklist.

Technical checklist

  • OK License identified Recorded: Apache-2.0
  • OK Maintenance signal Active
  • OK Verification status Official confirmed
  • OK Source evidence attached 1 source record(s)
  • NEEDS REVIEW Latest version recorded Not recorded

Official Links

Metadata & Governance

License Apache-2.0 — Inferred
Commercial modelUnknown
Maintenance statusActive
Verification status Official confirmed — Confirmed via the official repository API responses in SourceRefs below.
Latest versionNot recorded
Latest releaseNot recorded
Last activity2026-08-13
Last checked2026-08-13
First seenNot recorded

Dataset facts

Facts above come from the official dataset card only; unconfirmed fields stay unknown.

Related Resources & Dependencies

  • yolox — alternative to (verified)
  • detr — alternative to (verified)

Recent Activity

Related Knowledge

Collections