Voice and chat agent testing and monitoring
Hamming AI
Hamming AI builds a testing and monitoring platform for voice and chat agents, serving engineering teams and enterprises deploying conversational AI in healthcare, financial services and other industries. Teams can generate scenarios from an agent’s prompt, simulate calls with accents, background noise and interruptions, and check whether tasks such as bookings actually complete. Adversarial tests probe prompt injection, policy violations and personal-data leakage; production failures can become regression tests for future releases.
Founder and CEO Sumanyu Sharma previously led data work at Citizen and worked as a Senior Staff Data Scientist at Tesla. Hamming began with AI evaluation infrastructure after incorporation in 2023, then shifted its focus to voice-agent QA in 2024. Its engineering approach examines the audio pipeline alongside conversation outcomes: turn-by-turn diagnostics capture timing and transcription, separate latency across speech recognition, language models and speech synthesis, and identify mishandled interruptions that may look normal in a transcript.
As of February 2026, the company reported processing more than 10 million minutes across over 10,000 voice agents in its QA workflows. Its Cisco partnership, announced in 2025, lets Webex AI Agent customers add Hamming’s testing, monitoring and red-teaming to their existing deployments. Hamming also offers SOC 2 Type II controls and business associate agreements for healthcare workflows involving protected health information.
Explore the recordings
Hamming’s supplied archive contains one recording. Use it as a guide to voice-agent failure modes, reliability workflows, and adversarial testing. The claims below describe what the speaker reported at the time of recording; they do not establish today’s company leadership, product capabilities, or performance.
Start with the gap between fluent speech and completed actions
In The Crisis in Voice AI, Sumanyu Sharma explains how confident speech can conceal incorrect information or an action that never happened. His appointment-booking anecdote provides an entry point for examining whether an agent actually completes its task. Continue with examples of repetition, skipped verification, unauthorized discounts, and incomplete bookings. His scale and error-rate estimates are speaker-reported figures without a supplied measurement methodology.
Follow the reliability loop from discovery to verification
For an operational path through the recording, focus on Sharma’s sequence: identify failures, prioritize their frequency and severity, investigate and apply fixes, check effectiveness and regressions, and keep monitoring production. He distinguishes manual listening, rubric scoring, and cross-conversation analysis. His testing advice emphasizes varied wording, accents, styles, and combinations of intent, alongside live A/B tests for effects synthetic testing cannot establish.
Explore verification bypass and adversarial testing
Read the talk’s adversarial-testing discussion alongside its examples of skipped verification and unauthorized actions. Sharma reports verification bypass and unauthorized data disclosure through prompt injection, and says Hamming’s testing could break roughly one in five tested agents. Treat that as a recorded claim, not a current benchmark. His recommendations connect pre-deployment testing, monitoring of both callers and agents, and continuous red teaming where failures carry substantial costs.
1 talk
Newest first1 speaker at AIE
Affiliations reflect their AIE appearances, not necessarily current employment.
