Dippu Kumar Singh is a senior director of emerging data and analytics at Fujitsu North America, developing production AI systems for demanding enterprise environments. His work spans contact-center voice intelligence, computer vision, healthcare decision support, and the governance required when sensitive data drives consequential decisions.
Before focusing on generative AI, Singh designed large-scale data platforms and streaming pipelines using Spark, Kafka, and Hadoop, and developed data architectures for manufacturing and public-sector organizations. His work also includes markerless motion analytics, protecting computer-vision systems against data poisoning, and contributions to Fujitsu’s AI ethics and compliance assessments.
Turning difficult signals into dependable decisions
VoiceOps for contact centers. Singh’s voice-intelligence architecture converts overlapping customer-service conversations into structured records of customer intent, agent actions, and resolutions. Multichannel capture, transcription, language-model extraction, and CRM synchronization target after-call work reduction while giving managers more consistent operational data.
Privacy before model processing. Separate customer and agent audio channels preserve speaker attribution; noise filtering, specialized vocabulary, and numerical normalization improve transcription. Early PII masking prevents sensitive details from reaching downstream models unnecessarily, although additional safeguards introduce latency and architectural complexity.
Structured outputs with human oversight. Few-shot prompts, predefined intent categories, transcript-grounding checks, and JSON schemas make generated information usable by business systems. Human-in-the-loop verification keeps operators responsible for correcting and approving records. Singh’s proposed extensions include explainable coaching, demand-based staffing forecasts, and early detection of abusive calls.
Dippu Kumar Singh walks through an audio-to-CRM pipeline that reduces after-call documentation while preserving speaker roles, structured outputs, and operator confirmation.