Can LLM Agents Price Competitively? A Dynamic Multi-Attribute Auction Benchmark for Agentic Commerce
06:00 · August 4, 2026 · arXiv cs.AI RSS

Agentic commerce is moving from concept to deployed infrastructure: payment networks, retailers, and AI platforms are setting the stage for agents to transact on behalf of merchants and consumers. Yet whether the LLMs behind these agents can price competently in real markets, where customer preferences are hidden, competitors adapt in real time, and demand can shift without warning, has not been systematically tested. We introduce Bazaar, a dynamic sealed-bid benchmark for multi-attribute auction under these conditions. Despite its dynamics, the benchmark is grounded in closed-form customer utilities, enabling exact evaluation. Across 11 frontier LLMs from four providers, the leading agents on customer acquisition (e.g. Gemini 3.1 Pro) are often not the leading agents on profit (e.g. Opus 4.6). The ranking shifts again under demand shocks: agents that learned fastest pre-shock are typically the slowest to revise their beliefs afterwards, while Gemini 3.1 Pro recovers fastest despite not leading on profit. However, even the strongest agent captures less than a third of hindsight-optimal profit, suggesting current LLMs are progressing in agentic commerce but leave substantial headroom.
Summary
Bazaar is a benchmark that places LLM-based merchant agents inside repeated sealed-bid multi-attribute auctions. Each round, agents must configure a product across several attributes and set a price for customers whose valuations remain hidden. Only sparse feedback—winner identity, chosen configuration, and realized profit for the winner—is returned, forcing agents to infer customer preferences, manage margins, and compete against rivals whose costs differ. The environment also inserts unannounced preference shifts midway through each run, creating a controlled test of both initial learning and subsequent belief revision.
Evaluation across eleven frontier models from four providers shows that performance splits along two axes. Gemini 3.1 Pro achieves the highest win rate, while Opus 4.6 leads in total profit through stricter margin discipline. These rankings invert under demand shocks: models that climbed fastest before a shift tend to revise their strategies most slowly afterward, whereas Gemini 3.1 Pro recovers quickest despite not topping the profit table. Even the strongest agent, however, captures less than one-third of the profit attainable by an oracle with full knowledge of customer values.
The benchmark further isolates two recurring failure modes. Agents either under-configure offerings and miss profitable sales, or they win frequently yet leave margin on the table. Increasing the model’s thinking budget moves behavior along this surface, reducing missed wins at the cost of lower margins per transaction. The results indicate that current LLMs can acquire basic pricing competence in competitive, non-stationary markets, yet still exhibit a substantial gap relative to hindsight-optimal policies.
Why it matters
This research is highly relevant for Dutch AI researchers and fintech/e-commerce enterprises developing agentic commerce solutions. It provides a rigorous, reproducible benchmark for evaluating algorithmic pricing, margin discipline, and market adaptation in LLMs.





