WEKA Maximizes Token Output With Lower Cost Per Token on NVIDIA BlueField-4 STX - Thailand PR News

 - [Banking](https://thailandbusinessnews.net/category/banking/)
- [Companies](https://thailandbusinessnews.net/category/companies/)
- [Economy](https://thailandbusinessnews.net/category/economy/)
- [Press Releases](https://thailandbusinessnews.net/category/press-release/)
    - [PR Newswire](https://thailandbusinessnews.net/pr-newswire-updates/)
- [Tech](https://thailandbusinessnews.net/category/tech/)
- [Tourism](https://thailandbusinessnews.net/category/tourism/)

    - [Press Release](https://thailandbusinessnews.net/category/press-release/)

# WEKA Maximizes Token Output With Lower Cost Per Token on NVIDIA BlueField-4 STX

   #####  Up next

  ![LX Pantos Accelerates Global Expansion with Acquisition of Logistics Center in Poland](https://i0.wp.com/mma.prnasia.com/media2/2934363/image3.jpg?resize=200%2C110&ssl=1 "LX Pantos Accelerates Global Expansion with Acquisition of Logistics Center in Poland")

 [](https://thailandbusinessnews.net/press-release/lx-pantos-accelerates-global-expansion-with-acquisition-of-logistics-center-in-poland/)

 ###### [LX Pantos Accelerates Global Expansion with Acquisition of Logistics Center in Poland](https://thailandbusinessnews.net/press-release/lx-pantos-accelerates-global-expansion-with-acquisition-of-logistics-center-in-poland/)

  Published on 17 March 2026 #####  Author

#####   [ CISION PRNewswire ](https://thailandbusinessnews.net/author/cision/)

  ##### Tags

- [PRNewswire](https://thailandbusinessnews.net/tag/prnewswire/)

*NeuralMesh and Augmented Memory Grid Integration with NVIDIA STX Increases Token Production by 6.5x in the Same GPU Footprint, Slashing Cost of Inference for AI-Driven Organizations*

SAN JOSE, Calif. and CAMPBELL, Calif., March 17, 2026 /PRNewswire/ — From GTC 2026: [WEKA](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=1676554663&u=https%3A%2F%2Fwww.weka.io%2F&a=WEKA), the AI storage and memory systems company, today announced the integration of its [NeuralMesh](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=3591769974&u=https%3A%2F%2Fwww.weka.io%2Fproduct%2Fhow-it-works%2F&a=NeuralMesh)™ software with the [NVIDIA STX](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=2852396885&u=https%3A%2F%2Fwww.nvidia.com%2Fen-us%2Fdata-center%2Fai-storage%2Fstx%2F&a=NVIDIA+STX) reference architecture. WEKA’s breakthrough [Augmented Memory Grid](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=1438818156&u=https%3A%2F%2Fwww.weka.io%2Fproduct%2Faugmented-memory-grid%2F&a=Augmented+Memory+Grid)™ memory extension technology running on NeuralMesh will support NVIDIA STX to bring high-throughput context memory storage to agentic AI factories, making long-context reasoning seamless across sessions, tools, and tasks. Leveraging NVIDIA Vera Rubin NVL72, [NVIDIA BlueField-4](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=1671419341&u=https%3A%2F%2Fwww.nvidia.com%2Fen-us%2Fnetworking%2Fproducts%2Fdata-processing-unit%2F&a=NVIDIA+BlueField-4), and [NVIDIA Spectrum-X](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=797274279&u=https%3A%2F%2Fwww.nvidia.com%2Fen-us%2Fnetworking%2Fspectrumx%2F&a=NVIDIA+Spectrum-X) Ethernet, the NeuralMesh solution based on NVIDIA STX will deliver an estimated increase of 4-10x more tokens per second for context memory while supporting at least 320 GB read and 150 GB write throughput per second for AI workloads, more than double the throughput of conventional AI storage platforms.

 [![WEKA and NVIDIA unlock cost-efficient AI inference at scale](https://mma.prnasia.com/media2/2934399/WEKA_and_NVIDIA.jpg?p=medium600 "WEKA and NVIDIA unlock cost-efficient AI inference at scale")](https://mma.prnasia.com/media2/2934399/WEKA_and_NVIDIA.jpg?p=medium600)
 WEKA and NVIDIA unlock cost-efficient AI inference at scale

**Solving the Inference Cost Problem with Shared KV Cache Infrastructure**
Scaling agentic systems, especially for software engineering applications, exposes a hard truth: today’s AI economics are decided at the memory infrastructure layer. Every large-scale inference fleet hits the memory wall: limited high-bandwidth memory (HBM) on the GPU is rapidly exhausted, key-value (KV) cache is evicted, context is lost, and the system is forced to repeat work it already completed. This architectural inefficiency sends inference costs soaring. The answer is a shared KV cache infrastructure that keeps context live across agents, users, and sessions. It eliminates redundant computation, sustains token throughput, and maintains predictable performance. Without shared KV cache infrastructure, every increase in concurrent users and agents becomes a liability — costs rise, experiences degrade, and the inference fleet becomes harder to operate the larger it grows. With STX for context memory, NVIDIA is introducing a blueprint to address these core inference bottlenecks.

**Context Memory Storage: The Foundation of Agentic AI Factories**
With co-designed WEKA solutions based on NVIDIA STX architecture, AI clouds, enterprises, and AI model builders can deploy the infrastructure foundation they need to run GPUs at peak productivity, sustain high-volume token production, and make large-scale inference more energy and cost-efficient.

Leading AI innovators and cloud providers, such as [Firmus](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=1759237850&u=https%3A%2F%2Ffirmus.co%2F&a=Firmus), are already transforming their inference economics with Augmented Memory Grid on NeuralMesh.

"Real-world AI doesn’t run in a lab— it has power constraints, cooling limits, and relentless workload demand. Firmus is built for exactly that. Paired with NVIDIA AI infrastructure, WEKA Augmented Memory Grid delivers up to 6.5x higher tokens per second and 4x faster TTFT at scale, proving we can get more performance from the same GPU footprint. With NeuralMesh and Augmented Memory Grid integrated into our NVIDIA-aligned AI Factory and NVIDIA STX reference architecture, we’ll be able to deliver the fastest context memory network for predictable and efficient inference at scale," said Daniel Kearney, Chief Technology Officer at[ ](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=3606729127&u=https%3A%2F%2Ffirmus.co%2F&a=%C2%A0)Firmus.

**NeuralMesh and NVIDIA STX: Purpose-Built for Agentic AI**
NeuralMesh is WEKA’s intelligent, adaptive storage system built on over 170 patents. It will run across the full-stack STX reference architecture, providing the next-generation storage organizations need to standardize high-performance AI data services and accelerate agentic AI outcomes. WEKA’s Augmented Memory Grid is a purpose-built memory extension layer that pools and persists KV cache outside of GPU memory, keeping long-context sessions stable and concurrency high as inference workloads grow. First unveiled at GTC 2025 and generally available to NeuralMesh customers today, Augmented Memory Grid has been validated with Supermicro on [NVIDIA Grace](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=888123797&u=https%3A%2F%2Fwww.nvidia.com%2Fen-us%2Fdata-center%2Fgrace-cpu-superchip%2F&a=NVIDIA+Grace) CPUs and BlueField-3 DPUs to deliver numerous benefits that improve AI economics, including:

- **Faster User Experiences:** Augmented Memory Grid on NeuralMesh delivers up to 4-20x improvement in time-to-first-token, keeping AI agents and applications responsive under real-world load.
- **More Revenue from the Same Hardware:** Serve 6.5x more tokens per GPU — without adding infrastructure.
- **Sustained Performance at Scale:** Augmented Memory Grid maintains high KV cache hit rates even as sessions, agents, and context windows grow — preventing the performance cliff that hits DRAM-only architectures.
- **GPU-Native Efficiency:** BlueField-4 integration offloads the storage data path from the CPU, keeping GPUs fully productive and eliminating I/O bottlenecks.

"With coding LLMs advancing, we’re seeing unprecedented adoption of Agentic AI use cases for software engineering, where productivity increases by 100-1000x. As coding assistants make repeated calls against largely unchanged codebases and prompts, WEKA’s Augmented Memory Grid reuses cached context instead of forcing redundant prefill, even as context windows grow to incredible lengths. This provides a major boost in response times and greatly increases the number of concurrent users running on the same infrastructure," said Liran Zvibel, co-founder and CEO at WEKA. "WEKA first identified this need for context memory storage more than a year ago and launched Augmented Memory Grid at GTC 2025. Now, NVIDIA STX opens the door to organizations running their storage and memory extension infrastructure on state-of-the-art NVIDIA Vera Rubin architecture, including NVIDIA BlueField-4 and NVIDIA Spectrum-X Ethernet. Running Augmented Memory Grid on NeuralMesh for NVIDIA STX delivers extreme performance and efficiency that translates directly to game-changing AI economics."

**Availability**

WEKA’s Augmented Memory Grid is commercially available with NeuralMesh today.

Organizations that don’t address the memory wall today will find it harder and more expensive to scale tomorrow. As agentic workloads grow and context windows expand, DRAM-only architectures face a compounding cost problem: each additional concurrent user or session increases recomputation overhead, GPU idle time, and operational cost. The organizations that architect for persistent KV cache now will have a structural cost and performance advantage over those that wait.

For more information about NeuralMesh, visit: [weka.io/NeuralMesh](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=2989255427&u=https%3A%2F%2Fwww.weka.io%2Fneuralmesh%2F&a=weka.io%2FNeuralMesh).
For more information about Augmented Memory Grid, visit: [weka.io/augmented-memory-grid](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=2530118059&u=https%3A%2F%2Fwww.weka.io%2Faugmented-memory-grid%2F&a=weka.io%2Faugmented-memory-grid).

Organizations can learn more at [weka.io/nvidia](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=2995353283&u=https%3A%2F%2Fwww.weka.io%2Fnvidia&a=weka.io%2Fnvidia) or visit WEKA at GTC 2026, booth #1034.

**About WEKA**
WEKA is transforming how organizations build, run, and scale AI workflows with NeuralMesh™ by WEKA®, its intelligent, adaptive mesh storage system. Unlike traditional data infrastructure, which becomes slower and more fragile as workloads expand, NeuralMesh becomes faster, stronger, and more efficient as it scales, dynamically adapting to AI environments to provide a flexible foundation for enterprise AI and agentic AI innovation. Trusted by 30% of the Fortune 50, NeuralMesh helps leading enterprises, AI cloud providers, and AI builders optimize GPUs, scale AI faster, and reduce innovation costs. Learn more at [www.weka.io](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=1913752326&u=https%3A%2F%2Fwww.weka.io%2F&a=www.weka.io) or connect with us on [LinkedIn](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=2818042447&u=https%3A%2F%2Fwww.linkedin.com%2Fcompany%2Fweka-io&a=LinkedIn) and [X](https://edge.prnewswire.com/c/link/?t=0&l=en&o=4641359-1&h=3826677041&u=https%3A%2F%2Fx.com%2FWEKA&a=X).

*WEKA and the W logo are registered trademarks of WekaIO, Inc. Other trade names herein may be trademarks of their respective owners.*

 [![WEKA_v1_Logo_new](https://mma.prnasia.com/media2/1796062/WEKA_v1_Logo_new.jpg?p=medium600 "WEKA_v1_Logo_new")](https://mma.prnasia.com/media2/1796062/WEKA_v1_Logo_new.jpg?p=medium600)
 WEKA\_v1\_Logo\_new

[Source link ](http://www.prnasia.com/story/archive/4910377_AE10377_0?rand=89749)

```
This content was prepared by our news partner, Cision PR Newswire. The opinions and the content published on this page are the author’s own and do not necessarily reflect the views of Siam News Network
```

  #####  You May Also Like

  [](https://thailandbusinessnews.net/press-release/v-gallant-launches-malaysias-first-gpu-powered-ai-workspace-and-intelli-x-expands-ai-enablement-in-malaysia/) ![V Gallant Launches Malaysia's First GPU-Powered AI Workspace and Intelli-X, Expands AI Enablement in Malaysia](https://i0.wp.com/mma.prnasia.com/media2/2924505/The_Launching_of_V_Gallant_GPU_Lounge.jpg?resize=260%2C140&ssl=1 "V Gallant Launches Malaysia's First GPU-Powered AI Workspace and Intelli-X, Expands AI Enablement in Malaysia")

- [Press Release](https://thailandbusinessnews.net/category/press-release/)

 ## V Gallant Launches Malaysia's First GPU-Powered AI Workspace and Intelli-X, Expands AI Enablement in Malaysia

  Grand launch of the V Gallant GPU Lounge and beta testing for Intelli-X,…

- [CISION PRNewswire](https://thailandbusinessnews.net/author/cision/ "View all posts by CISION PRNewswire")
- March 3, 2026

   [](https://thailandbusinessnews.net/press-release/hyundai-mobis-debuts-holographic-heads-up-display-redefining-in-car-tech-at-ces-2025/) ![Hyundai Mobis Debuts Holographic Heads-Up Display, Redefining In-Car Tech at CES 2025](https://i0.wp.com/r2.thailandbusinessnews.net/2025/01/Hyundai_Mobis_and_ZEISS.jpg?resize=260%2C140&ssl=1 "Hyundai Mobis Debuts Holographic Heads-Up Display, Redefining In-Car Tech at CES 2025")

- [Press Release](https://thailandbusinessnews.net/category/press-release/)

 ## Hyundai Mobis Debuts Holographic Heads-Up Display, Redefining In-Car Tech at CES 2025

  Hyundai Mobis revealed its Holographic Windshield Display technology, which projects driving information…

- [CISION PRNewswire](https://thailandbusinessnews.net/author/cision/ "View all posts by CISION PRNewswire")
- January 9, 2025

   [](https://thailandbusinessnews.net/press-release/accelerate-ai-native-amplify-success-huawei-cloud-unveiled-new-cloud-services-and-solutions-at-mwc-2025/) ![Accelerate AI-Native, Amplify Success: Huawei Cloud Unveiled New Cloud Services and Solutions at MWC 2025](https://i0.wp.com/r2.thailandbusinessnews.net/2025/03/Jacqueline_Shi.jpg?resize=260%2C140&ssl=1 "Accelerate AI-Native, Amplify Success: Huawei Cloud Unveiled New Cloud Services and Solutions at MWC 2025")

- [Press Release](https://thailandbusinessnews.net/category/press-release/)

 ## Accelerate AI-Native, Amplify Success: Huawei Cloud Unveiled New Cloud Services and Solutions at MWC 2025

  BARCELONA, Spain, March 5, 2025 /PRNewswire/ — The Huawei Cloud Summit, with the…

- [CISION PRNewswire](https://thailandbusinessnews.net/author/cision/ "View all posts by CISION PRNewswire")
- March 4, 2025

 ##### Most Read

- [![HiBy Officially Unveils Cyberpunk 2077 Collaboration Products](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/hiby-officially-unveils-cyberpunk-2077-collaboration-products/ "HiBy Officially Unveils Cyberpunk 2077 Collaboration Products") [HiBy Officially Unveils Cyberpunk 2077 Collaboration Products](https://thailandbusinessnews.net/press-release/hiby-officially-unveils-cyberpunk-2077-collaboration-products/ "HiBy Officially Unveils Cyberpunk 2077 Collaboration Products")
- [![Nel ASA: Receives purchase order from Collins Aerospace for US Navy stacks](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/nel-asa-receives-purchase-order-from-collins-aerospace-for-us-navy-stacks/ "Nel ASA: Receives purchase order from Collins Aerospace for US Navy stacks") [Nel ASA: Receives purchase order from Collins Aerospace for US Navy stacks](https://thailandbusinessnews.net/press-release/nel-asa-receives-purchase-order-from-collins-aerospace-for-us-navy-stacks/ "Nel ASA: Receives purchase order from Collins Aerospace for US Navy stacks")
- [![Sumsub Agentic AI Council Launches in APAC and Invites Experts to Advance Trusted AI Adoption](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/sumsub-agentic-ai-council-launches-in-apac-and-invites-experts-to-advance-trusted-ai-adoption/ "Sumsub Agentic AI Council Launches in APAC and Invites Experts to Advance Trusted AI Adoption") [Sumsub Agentic AI Council Launches in APAC and Invites Experts to Advance Trusted AI Adoption](https://thailandbusinessnews.net/press-release/sumsub-agentic-ai-council-launches-in-apac-and-invites-experts-to-advance-trusted-ai-adoption/ "Sumsub Agentic AI Council Launches in APAC and Invites Experts to Advance Trusted AI Adoption")
- [![More Than a Stylus: How Hands-On Experiences Are Changing the Way Consumer Electronics Are Presented](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/more-than-a-stylus-how-hands-on-experiences-are-changing-the-way-consumer-electronics-are-presented/ "More Than a Stylus: How Hands-On Experiences Are Changing the Way Consumer Electronics Are Presented") [More Than a Stylus: How Hands-On Experiences Are Changing the Way Consumer Electronics Are Presented](https://thailandbusinessnews.net/press-release/more-than-a-stylus-how-hands-on-experiences-are-changing-the-way-consumer-electronics-are-presented/ "More Than a Stylus: How Hands-On Experiences Are Changing the Way Consumer Electronics Are Presented")
- [![Yuyu Teijin Medicare Uses Drones to Deliver Medical Supplies to Patients on Korea's Remote Islands](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/yuyu-teijin-medicare-uses-drones-to-deliver-medical-supplies-to-patients-on-koreas-remote-islands/ "Yuyu Teijin Medicare Uses Drones to Deliver Medical Supplies to Patients on Korea's Remote Islands") [Yuyu Teijin Medicare Uses Drones to Deliver Medical Supplies to Patients on Korea's Remote Islands](https://thailandbusinessnews.net/press-release/yuyu-teijin-medicare-uses-drones-to-deliver-medical-supplies-to-patients-on-koreas-remote-islands/ "Yuyu Teijin Medicare Uses Drones to Deliver Medical Supplies to Patients on Korea's Remote Islands")
- [![Dragonpass: AI Is Shifting the Membership Economy from Perks Redemption to Personalized Experiences](https://secure.gravatar.com/avatar/96e712b5408b51145de772098c3ff3417ce8a11555b7c388d84204c5769f5631?s=40&d=blank&r=g)](https://thailandbusinessnews.net/press-release/dragonpass-ai-is-shifting-the-membership-economy-from-perks-redemption-to-personalized-experiences/ "Dragonpass: AI Is Shifting the Membership Economy from Perks Redemption to Personalized Experiences") [Dragonpass: AI Is Shifting the Membership Economy from Perks Redemption to Personalized Experiences](https://thailandbusinessnews.net/press-release/dragonpass-ai-is-shifting-the-membership-economy-from-perks-redemption-to-personalized-experiences/ "Dragonpass: AI Is Shifting the Membership Economy from Perks Redemption to Personalized Experiences")

##### Subscribe via Email

 Enter your email address to subscribe and receive notifications of new posts by email.

  Email Address

       Subscribe

  [](https://thailandbusinessnews.net/press-release/almost-95-of-viral-korean-skincare-and-glass-skin-tiktok-videos-contain-misleading-claims-new-study-shows/) ## Almost 95% of viral Korean skincare and 'glass skin' TikTok videos contain misleading claims, new study shows

  VIENNA, Sept. 30, 2026 /PRNewswire/ — Almost 95% of viral Korean skincare…

   [](https://thailandbusinessnews.net/press-release/su-group-narrows-first-half-operating-loss-executes-on-long-term-strategy/) ## SU Group Narrows First-Half Operating Loss, Executes on Long-Term Strategy

  Business Momentum Led by Public-Sector Awards Across Healthcare, Cultural and Border…

   [](https://thailandbusinessnews.net/press-release/frost-sullivan-welcomes-back-sarwant-singh-to-accelerate-digital-innovation-and-global-growth/) ## Frost &amp; Sullivan Welcomes Back Sarwant Singh to Accelerate Digital Innovation and Global Growth

  Industry veteran and futurist returns to advance digital platforms, strengthen advisory…

   [](https://thailandbusinessnews.net/press-release/ontario-contributes-1-7-million-investment-to-advance-life-sciences-innovation-at-piramal-pharmas-aurora-facility/) ![Ontario Contributes .7 Million Investment to Advance Life Sciences Innovation at Piramal Pharma's Aurora Facility](https://i0.wp.com/mmx.prnasia.com/media/MS1998885/Peter-DeYoung-CEO-Piramal-Global-Pharma.jpg?resize=260%2C140&ssl=1 "Ontario Contributes .7 Million Investment to Advance Life Sciences Innovation at Piramal Pharma's Aurora Facility")

 ## Ontario Contributes $1.7 Million Investment to Advance Life Sciences Innovation at Piramal Pharma's Aurora Facility

  Piramal Pharma Solutions is investing over $5.3 million in its Aurora drug…
<!-- Performance optimized by Docket Cache: https://wordpress.org/plugins/docket-cache -->
