← All speakers

Bio, Work & Ideas

Arjun Bansal

Conference affiliation: Log10 · 2024

On this page

Arjun Bansal is the co-founder and chief executive of Everest by Log10, which develops artificial-intelligence systems for clinical documentation, regulatory submissions, and other demanding life-sciences workflows. Previously, he co-founded Nervana Systems, the deep-learning company Intel acquired in 2016, and led AI software and research at Intel.

From neuroscience to enterprise AI

Bansal studied computer science at Caltech, earned a neuroscience doctorate at Brown University, and completed postdoctoral training at Boston Children’s Hospital and Harvard Medical School. His research combined machine learning with neurophysiological data collection and analysis, grounding his technical career in both computational methods and biological experimentation.

At Nervana Systems, Bansal led algorithms, machine-learning software, and data science as the company developed an integrated deep-learning hardware and software stack. Following its acquisition, he became an Intel vice president overseeing AI software and research, including healthcare partnerships with GE, Siemens, and Philips.

He subsequently co-founded and led XOKind, whose Una travel-planning assistant combined collaborative planning with personalized recommendations. He later co-founded Log10 to tackle a fundamental obstacle to deploying AI applications: measuring whether their outputs actually meet professional standards.

Teaching AI systems to recognize quality

Bansal’s work centers on human-calibrated automated evaluation: preserving expert judgment while making quality assessment fast enough for production systems. His approach has several distinctive components:

  • Reviewer-specific evaluation models. Generic AI judges can favor longer responses, their own outputs, or whichever answer appears first. In research co-authored with Ansup Babu, Bansal developed evaluators trained against explicit grading criteria and individual reviewers’ judgments; reported summarization experiments reduced absolute error by approximately 44 percent.
  • Synthetic bootstrapping from scarce expert feedback. Those experiments also used small sets of human-reviewed examples to generate additional training data, allowing evaluators initialized with roughly 25 to 50 expert labels to approach the performance of systems trained on substantially larger labeled datasets.
  • AutoFeedback as production infrastructure. Log10’s open-source integration and evaluation tooling connects application traces, grading rubrics, automated assessments, and model comparisons. At AI Engineer World’s Fair 2024, Bansal explained how these feedback loops can detect hallucinations, prioritize human review, curate training data, and improve prompts or fine-tuned models.

Through Everest, Bansal now applies this evaluation-first approach to clinical-study reports, safety documentation, and regulatory submissions—workflows where specialized expertise, traceability, and reliable review directly determine whether generated material is usable.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Echo AI’s customer-conversation pipeline shows how broad analysis becomes useful only when teams can inspect outputs, grade them against customer needs, and turn corrections into better models.

  • Customer understanding breaks down at scale
    0:00 ↗
  • From sampling to discovery
    2:18 ↗
  • One message can expose several problems
    3:58 ↗
  • The pipeline starts before the prompt
    5:29 ↗
  • What an accuracy problem looks like in production
    7:45 ↗
  • The evaluator needs evaluation too
    9:47 ↗
  • Build the evaluator around human judgments
    11:45 ↗
  • Put feedback beside logging and debugging
    13:52 ↗
  • A broken TV and a network of downstream analyses
    14:58 ↗
  • Make a correction easy to capture
    17:10 ↗
  • Inspect a failed summary, not just its score
    18:25 ↗
  • From feedback to a reported customer gain
    19:22 ↗
  • Carry the workflow into the application
    20:15 ↗

References