Methodology
Loopbench measures email platforms the way agentic operators feel them: can you authenticate, send quickly, stay reliable, produce on-brand creative, express automations, and let an agent drive the loop? Latest protocol window: 2026-06-02 to 2026-06-20.
Window facts
- Dates
- 2026-06-02 to 2026-06-20
- Providers
- 12
- Lab region
- EU + US dual runners
- Publisher
- Loopbench Independent Testing Lab
Metric definitions
Deliverability
Composite of authentication readiness (SPF, DKIM, DMARC tooling), seed-list inbox placement in our controlled panels, and bounce hygiene signals during the test window.
Send latency
Median wall-clock time from authenticated API accept to provider-acknowledged queue for a 1-recipient transactional send in our lab region.
Reliability
Success rate of scheduled and triggered sends across the window, weighted by recovery after injected soft failures in our harness.
Agent surface
How completely an external agent can create, schedule, send, and read outcomes via API and/or MCP without a human in the dashboard.
Design quality
Blind review of on-brand visual fidelity and cross-client rendering for emails produced under a fixed creative brief.
Automation depth
Ability to express welcome, nurture, and event-triggered flows, including branching and publish reliability from API or UI.
Overall composite
Overall = 0.2 deliverability + 0.15 reliability + 0.1 latencyScore + 0.2 agentSurface + 0.2 designQuality + 0.15 automationDepth. latencyScore maps p50 milliseconds onto a 0-10 scale capped at 1000ms. Sending shortlists on /compare use a dedicated sending composite (0.4 deliverability + 0.35 reliability + 0.25 latencyScore).
Sample sizes
Each provider page lists its send sample size for the window. Seed-panel placement uses a smaller fixed panel shared across providers. We publish n so readers can weight confidence.
What we refuse to invent
We do not fabricate funding rounds, customer counts, revenue lifts, or quotes from real people. Qualitative positioning comes from vendor docs and public product pages, linked inline.
External references we align with
Authentication and bulk-sender expectations are grounded in Gmail sender guidelines, Yahoo sender requirements, M3AAWG published documents, DMARC (RFC 7489), and SMTP (RFC 5321). Rendering context comes from Litmus, Email on Acid, Really Good Emails, and Can I email.
FAQ
Are these inbox placement guarantees?
No. Seed panels are controlled and small relative to production audiences. Treat scores as comparative lab composites.
Why include agent surface?
Agentic loops fail when creative and send require constant dashboard work. We score whether an external agent can create, send, and read outcomes via API or MCP.
How often do you revise scores?
Each dated window can revise composites. We keep sparklines so directionality is visible without inventing fake precision about customer results.