QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

Wenfang Sun; Yingjun Du; Gaowen Liu; Yefeng Zheng; Cees G. M. Snoek

arXiv:2411.19534·cs.CV·December 16, 2025

QUOTA: Quantifying Objects with Text-to-Image Models for Any Domain

Wenfang Sun, Yingjun Du, Gaowen Liu, Yefeng Zheng, Cees G. M. Snoek

PDF

Open Access

TL;DR

QUOTA introduces a domain-agnostic, optimization-based framework for accurate object quantification in text-to-image models without retraining, effectively handling unseen domains and stylistic variations.

Contribution

It presents the first domain-agnostic, meta-learning approach for object quantification in text-to-image models, avoiding retraining and improving scalability.

Findings

01

Outperforms existing models in quantification accuracy

02

Maintains high accuracy across unseen domains

03

Sets new benchmarks in domain generalization for object counting

Abstract

We tackle the problem of quantifying the number of objects by a generative text-to-image model. Rather than retraining such a model for each new image domain of interest, which leads to high computational costs and limited scalability, we are the first to consider this problem from a domain-agnostic perspective. We propose QUOTA, an optimization framework for text-to-image models that enables effective object quantification across unseen domains without retraining. It leverages a dual-loop meta-learning strategy to optimize a domain-invariant prompt. Further, by integrating prompt learning with learnable counting and domain tokens, our method captures stylistic variations and maintains accuracy, even for object classes not encountered during training. For evaluation, we adopt a new benchmark specifically designed for object quantification in domain generalization, enabling rigorous…

Peer Reviews

No public reviews on file for this paper yet. If you reviewed it on a platform where reviews are public (OpenReview, ICLR, NeurIPS, ICML), you can paste yours below so the community can read it here.

Videos

No videos yet. Explain this paper in a talk, walkthrough, or lecture? Add one.

Taxonomy

TopicsImage Retrieval and Classification Techniques · Handwritten Text Recognition Techniques · Multimodal Machine Learning Applications

MethodsADaptive gradient method with the OPTimal convergence rate