# AI/ML

**URL:** https://community.amperecomputing.com/c/ai-ml/12.md

[Latest](https://community.amperecomputing.com/latest.md) · [Categories](https://community.amperecomputing.com/categories.md) · [Tags](https://community.amperecomputing.com/tags.md)

---

## [About the AI/ML category](https://community.amperecomputing.com/t/about-the-ai-ml-category/165)

<div class="topic-metadata">

**Author:** [@Aaron](https://community.amperecomputing.com/u/Aaron)\
**Replies:** 0

</div>

Welcome to the AI/ML category! This category is dedicated to all things related to artificial intelligence and machine learning. Here, you’ll find a wealth of resources and discussions about AI/ML, including news articl…

---

## [Curious If Anyone Has Real AI Production Experience With the B200 Yet](https://community.amperecomputing.com/t/curious-if-anyone-has-real-ai-production-experience-with-the-b200-yet/3449)

<div class="topic-metadata">

**Author:** [@brevisWillTalk](https://community.amperecomputing.com/u/brevisWillTalk)\
**Replies:** 1\
**Last updated:** [August 10, 2026, 11:08am UTC](https://community.amperecomputing.com/t/curious-if-anyone-has-real-ai-production-experience-with-the-b200-yet/3449 "2026-08-10T11:08:01Z")

</div>

I manage infrastructure for a small healthcare AI startup and we have been planning our next GPU investment for about two months now. Our current setup is handling the workload but barely and the roadmap we are committed…

---

## [CPU inference for Deepseek V4 flash?](https://community.amperecomputing.com/t/cpu-inference-for-deepseek-v4-flash/3434)

<div class="topic-metadata">

**Author:** [@dylanetaft](https://community.amperecomputing.com/u/dylanetaft)\
**Replies:** 5\
**Last updated:** [July 11, 2026, 6:09pm UTC](https://community.amperecomputing.com/t/cpu-inference-for-deepseek-v4-flash/3434 "2026-07-11T18:09:47Z")

</div>

Has anyone tested performance of deepseek v4 flash on AmpereOne, on CPU, on llama.cpp? I am wondering how some of the larger MoE models perform. I use Qwen 3.6 35b a3b on Altra and it is very decent, but I don’t have e…

---

## [Hardware for local community college](https://community.amperecomputing.com/t/hardware-for-local-community-college/3396)

<div class="topic-metadata">

**Author:** [@dohertyctl](https://community.amperecomputing.com/u/dohertyctl)\
**Replies:** 3\
**Last updated:** [May 13, 2026, 4:08pm UTC](https://community.amperecomputing.com/t/hardware-for-local-community-college/3396 "2026-05-13T16:08:50Z")

</div>

I lend time at a local Community College that supports a large population in central Massachusetts. When speaking with some of the faculty about planned upgrades and implementations for the summer, I found that they are …

---

## [Qwen3.5-35B-A3B benchmarks on AmpereOne](https://community.amperecomputing.com/t/qwen3-5-35b-a3b-benchmarks-on-ampereone/3336)

<div class="topic-metadata">

**Author:** [@lu\_zero](https://community.amperecomputing.com/u/lu_zero)\
**Replies:** 25\
**Last updated:** [April 24, 2026, 5:03am UTC](https://community.amperecomputing.com/t/qwen3-5-35b-a3b-benchmarks-on-ampereone/3336 "2026-04-24T05:03:32Z")

</div>

I started testing a bit and I’m wondering if I’m doing something wrong: llama-bench -m .cache/llama.cpp/unsloth\_Qwen3.5-35B-A3B-GGUF\_Qwen3.5-35B-A3B-UD-Q4\_K\_XL.gguf -t 96 -p 2048 -n 256 -r 3 model size params b…

---

## [Fitting devstral-2, olmo-3 on AmpereOne/AltraMax](https://community.amperecomputing.com/t/fitting-devstral-2-olmo-3-on-ampereone-altramax/3286)

<div class="topic-metadata">

**Author:** [@lu\_zero](https://community.amperecomputing.com/u/lu_zero)\
**Replies:** 3\
**Last updated:** [February 25, 2026, 8:39am UTC](https://community.amperecomputing.com/t/fitting-devstral-2-olmo-3-on-ampereone-altramax/3286 "2026-02-25T08:39:30Z")

</div>

There are a bunch of completely/near-completely open models that are very interesting, but documented to work well on gpu. Can somebody with hardware access try them on the current ampere hardware so there is an additio…

---

## [Inference in ONNX in C?](https://community.amperecomputing.com/t/inference-in-onnx-in-c/3238)

<div class="topic-metadata">

**Author:** [@dylanetaft](https://community.amperecomputing.com/u/dylanetaft)\
**Replies:** 4\
**Last updated:** [October 15, 2025, 8:14pm UTC](https://community.amperecomputing.com/t/inference-in-onnx-in-c/3238 "2025-10-15T20:14:48Z")

</div>

I have been experimenting with OnnxRuntime in C, C++, as well as NCNN, and a Yolo v12 image model trained in Pytorch. One thing I am noticing - the out of the box optimizations don’t quite work for inference on Ampere C…

---

## [Rebuilding AI libraries for Ampere CPU](https://community.amperecomputing.com/t/rebuilding-ai-libraries-for-ampere-cpu/3155)

<div class="topic-metadata">

**Author:** [@quocbao](https://community.amperecomputing.com/u/quocbao)\
**Replies:** 4\
**Last updated:** [October 4, 2025, 2:52pm UTC](https://community.amperecomputing.com/t/rebuilding-ai-libraries-for-ampere-cpu/3155 "2025-10-04T14:52:05Z")

</div>

Just want to share with you guys some of my prebuilt wheel files for AI-stuff with CUDA support. My stack is: Ubuntu 24.04 Python 3.12 CUDA 12.8 GCC 13 Many of them only have release for x86. Some libs like PyTorch …

---

## [New quantization methods for llama.cpp](https://community.amperecomputing.com/t/new-quantization-methods-for-llama-cpp/3064)

<div class="topic-metadata">

**Author:** [@binh](https://community.amperecomputing.com/u/binh)\
**Replies:** 6\
**Last updated:** [August 19, 2025, 10:51pm UTC](https://community.amperecomputing.com/t/new-quantization-methods-for-llama-cpp/3064 "2025-08-19T22:51:21Z")

</div>

I noticed that Ampere optimized llama.cpp recommends using two new model quantization methods, Q4\_K\_4 and Q8R16. Has anyone tried both quantization types? If so, could you share how they performed?

---

## [Hosting and scaling LLMs on OKE for production-grade GenAI solutions](https://community.amperecomputing.com/t/hosting-and-scaling-llms-on-oke-for-production-grade-genai-solutions/1013)

<div class="topic-metadata">

**Author:** [@Aaron](https://community.amperecomputing.com/u/Aaron)\
**Replies:** 5\
**Last updated:** [December 6, 2024, 5:11pm UTC](https://community.amperecomputing.com/t/hosting-and-scaling-llms-on-oke-for-production-grade-genai-solutions/1013 "2024-12-06T17:11:02Z")

</div>

This is a nice blog post by OCI talking about best ways to run AI models on OCI. In this post, we explore an efficient, cost-effective approach to hosting and scaling LLMs, specifically Meta’s Llama 3 models, on OCI Kub…

---

## [AI News interviewing Victor Jakubiuk](https://community.amperecomputing.com/t/ai-news-interviewing-victor-jakubiuk/554)

<div class="topic-metadata">

**Author:** [@Aaron](https://community.amperecomputing.com/u/Aaron)\
**Replies:** 0\
**Last updated:** [December 6, 2023, 9:13pm UTC](https://community.amperecomputing.com/t/ai-news-interviewing-victor-jakubiuk/554 "2023-12-06T21:13:27Z")

</div>

AI News interviewed Victor Jakubiuk, Head of AI at Ampere. One of the topics the article brings up is amount of energy AI is going take up. A new study says that by 2027, it could take the same amount of energy as Swe…

---

## [Weekend Read - What is the best processor for AI?](https://community.amperecomputing.com/t/weekend-read-what-is-the-best-processor-for-ai/472)

<div class="topic-metadata">

**Author:** [@Aaron](https://community.amperecomputing.com/u/Aaron)\
**Replies:** 1\
**Last updated:** [October 4, 2023, 11:44am UTC](https://community.amperecomputing.com/t/weekend-read-what-is-the-best-processor-for-ai/472 "2023-10-04T11:44:11Z")

</div>

I remember early in my career working for the Retail group of a large company and I started learning about what was the “Best” price for a product. And the answer was always, “it depends, what is your goal? Sell more? …

---

## [Ampere AI Support Information](https://community.amperecomputing.com/t/ampere-ai-support-information/198)

<div class="topic-metadata">

**Author:** [@kkrysa](https://community.amperecomputing.com/u/kkrysa)\
**Replies:** 0\
**Last updated:** [February 27, 2023, 5:21pm UTC](https://community.amperecomputing.com/t/ampere-ai-support-information/198 "2023-02-27T17:21:27Z")

</div>

Hi all, I’ll be periodically posting links to some of our resources on Ampere Optimized Frameworks. You might also want to follow Ampere’s LinkedIn for any ongoing updates. I’ll try and include them here as well if ther…

---

## [Ampere AI + Matoha Case Study: AI Training on CPU Instances Alone](https://community.amperecomputing.com/t/ampere-ai-matoha-case-study-ai-training-on-cpu-instances-alone/119)

<div class="topic-metadata">

**Author:** [@kkrysa](https://community.amperecomputing.com/u/kkrysa)\
**Replies:** 2\
**Last updated:** [February 21, 2023, 5:57pm UTC](https://community.amperecomputing.com/t/ampere-ai-matoha-case-study-ai-training-on-cpu-instances-alone/119 "2023-02-21T17:57:37Z")

</div>

Ampere AI conducted a rather unique case study satisfying AI training performance needs with CPU instances alone. The use of OCI Ampere A1 instances decreased the time needed for training by 30%. You’ll find more detail…
