Skip to content
LoopbenchBoard
BenchmarksProvidersCompareIncidentsArticlesMethodAbout

Methodology

Loopbench measures email platforms the way agentic operators feel them: can you authenticate, send quickly, stay reliable, produce on-brand creative, express automations, and let an agent drive the loop? Latest protocol window: 2026-06-02 to 2026-06-20.

Protocol binder open on a measurement bench

Window facts

Dates
2026-06-02 to 2026-06-20
Providers
12
Lab region
EU + US dual runners
Publisher
Loopbench Independent Testing Lab

Metric definitions

Deliverability

Composite of authentication readiness (SPF, DKIM, DMARC tooling), seed-list inbox placement in our controlled panels, and bounce hygiene signals during the test window.

Send latency

Median wall-clock time from authenticated API accept to provider-acknowledged queue for a 1-recipient transactional send in our lab region.

Reliability

Success rate of scheduled and triggered sends across the window, weighted by recovery after injected soft failures in our harness.

Agent surface

How completely an external agent can create, schedule, send, and read outcomes via API and/or MCP without a human in the dashboard.

Design quality

Blind review of on-brand visual fidelity and cross-client rendering for emails produced under a fixed creative brief.

Automation depth

Ability to express welcome, nurture, and event-triggered flows, including branching and publish reliability from API or UI.

Overall composite

Overall = 0.2 deliverability + 0.15 reliability + 0.1 latencyScore + 0.2 agentSurface + 0.2 designQuality + 0.15 automationDepth. latencyScore maps p50 milliseconds onto a 0-10 scale capped at 1000ms. Sending shortlists on /compare use a dedicated sending composite (0.4 deliverability + 0.35 reliability + 0.25 latencyScore).

Sample sizes

Each provider page lists its send sample size for the window. Seed-panel placement uses a smaller fixed panel shared across providers. We publish n so readers can weight confidence.

What we refuse to invent

We do not fabricate funding rounds, customer counts, revenue lifts, or quotes from real people. Qualitative positioning comes from vendor docs and public product pages, linked inline.

External references we align with

Authentication and bulk-sender expectations are grounded in Gmail sender guidelines, Yahoo sender requirements, M3AAWG published documents, DMARC (RFC 7489), and SMTP (RFC 5321). Rendering context comes from Litmus, Email on Acid, Really Good Emails, and Can I email.

Seed list inbox panel used during deliverability tests

FAQ

Are these inbox placement guarantees?

No. Seed panels are controlled and small relative to production audiences. Treat scores as comparative lab composites.

Why include agent surface?

Agentic loops fail when creative and send require constant dashboard work. We score whether an external agent can create, send, and read outcomes via API or MCP.

How often do you revise scores?

Each dated window can revise composites. We keep sparklines so directionality is visible without inventing fake precision about customer results.