← All speakers

Bio, Work & Ideas

Val Bercovici

Conference affiliation: WEKA · 2025

Val Bercovici is chief AI officer at WEKA, focused on the infrastructure economics that determine whether AI agents can operate at scale. He argues that persistent context, cache efficiency, and memory architecture increasingly dictate the cost and responsiveness of production inference.

Bercovici held senior technical and cloud-strategy roles at NetApp, including chief technology officer-at-large, and subsequently became chief technology officer of NetApp/SolidFire. He chaired the Storage Networking Industry Association’s Cloud Storage Initiative, contributing to the Cloud Data Management Interface, and represented NetApp on the Cloud Native Computing Foundation governing board.

He later founded and led PencilDATA, applying blockchain to enterprise data integrity, and also led cybersecurity company Chainkit. At WEKA, he directs AI and open-source strategy, extending his background in storage and distributed infrastructure into the economics of agentic inference.

  • AI token economics: Bercovici links token costs and inference throughput directly to storage latency, context caching, and accelerator utilization. His public writing connects these efficiencies to practical AI deployment across business and science.
  • The AI memory wall: Persistent, multi-turn agents can outgrow high-bandwidth GPU memory, forcing systems to reconstruct context and spend additional compute. Bercovici treats this memory bottleneck as an architectural challenge requiring faster, larger memory tiers.
  • Context-platform engineering: At AI Engineer Code 2025, Bercovici and Callan Fox introduced an open-source toolkit for modeling agent swarms and translating application-level service agreements into infrastructure objectives. Bercovici articulated the platform strategy and token economics; Fox developed the load generator and presented WEKA Labs’ research and benchmarks.
  • KV-cache reuse and persistent context: Bercovici argues that short cache lifetimes and repeated prefills inflate inference costs and exhaust subscription limits. His analysis of AI unit economics connects these problems to Augmented Memory Grid, WEKA’s approach to extending inference memory through persistent, high-performance storage.

Read the topics behind these talks

1 conference talk

Key ideas

Scroll to read ↓

Agent context is reusable only if the inference platform retains it and retrieves it quickly enough. WEKA’s toolkit explores how cache lifetime, memory tiers and workload shape affect that reuse.

  • How do you keep an agent’s context available?
    0:09 ↗
  • The prediction problem behind prompt-cache pricing
    2:27 ↗
  • Two feedback loops, one storage obligation
    3:51 ↗
  • Inside a coding-agent conversation
    5:36 ↗
  • More agents create more opportunities—and more demand
    8:18 ↗
  • Longer retention trades memory for less prefill
    10:45 ↗
  • The provider needs a productive hit-rate band
    13:46 ↗
  • Token storage needs capacity and transfer speed
    16:05 ↗
  • Exercise the working set, not just the server
    18:51 ↗
  • What changes when the workload outgrows DRAM?
    20:26 ↗
  • Put the workload assumptions to the test
    23:07 ↗

References