Mengenai Kami

Perundingan yang lahir di persimpangan ekonomi tingkah laku dan pengalaman manusia.

Kini Mengambil Pekerja

Sertai pasukan yang membentuk semula cara dunia mengalami jenama.

Lihat peranan yang tersedia →

SYARIKAT

BERKEMBANG BERSAMA KAMI

HUBUNG

Perkhidmatan

Perundingan pengurusan dan CX komprehensif untuk jenama perusahaan.

SEMUA PERKHIDMATAN

Terokai rangkaian penuh perkhidmatan perundingan CX & pengurusan.

Lihat semua layanan →

TERAS

PAKAR

Penyelesaian

Penyelesaian berstruktur yang mengubah cita-cita CX menjadi hasil yang boleh diukur.

SEMUA PENYELESAIAN

Terokai setiap penyelesaian CX yang kami tawarkan.

Lihat penyelesaian →

STRATEGI & TADBIR URUS

REKA BENTUK & PENYAMPAIAN

BUDAYA & PENGALAMAN

Industri-industri

Satu dekad transformasi CX merentasi sektor-sektor utama di rantau ini.

SEMUA INDUSTRI

Lihat bagaimana kami bekerja merentasi setiap sektor.

Semak industri →

PERSEKITARAN BINAAN

KEWANGAN & TEKNOLOGI

ORANG & MOBILITI

Produk

Alat, platform, dan AI proprietari yang menggerakkan transformasi CX.

SEMUA PRODUK

Terokai ekosistem produk Renascence yang lengkap.

Lihat produk →

AI & TEKNOLOGI

PEMBELAJARAN & PERMAINAN

PLATFORM & ALATAN

PRODUK AI

Pendapat

Wawasan, penyelidikan dan perbualan di barisan hadapan CX.

BacaJurnal PengalamanArtikel & penyelidikan mengenai CX, tingkah laku dan transformasi.Tonton & DengarPengalaman LoomPodcast video kami tentang CX & tingkah laku.TersusunBerita CXBerita industri penting dalam CX, tanpa kebisingan.

Artikel terkini

Episod terkini

Berita terkini

Hab

Alat, templat, dan sumber percuma untuk memajukan amalan CX anda.

BAHARU · MANIFESTO

Bakar Dek. Sepuluh Kebaikan. Tiada Alasan. — baca manifesto kami untuk perunding yang berani.

Mula membaca →

ALAT AI

ALAT PERCUMA

PEMBELAJARAN

BUDAYA

AI · 19 September 2026

DeepSeek V4.1-Flash Cuts AI Agent Memory Costs by 75%

DeepSeek's new V4.1-Flash model shrinks key-value cache memory to about a quarter of its predecessor's, lowering the cost of running AI agents while matching or beating rivals on coding benchmarks.

Pusat Berita
Taklimat terpilih · 2 min bacaan

What happened

Chinese AI developer DeepSeek has released V4.1-Flash, a new multimodal model built to make AI agents significantly cheaper to run. The model uses a mixture-of-experts design with 552 billion total parameters, but activates only 16 billion per token, and it cuts the memory needed for its key-value cache to roughly a quarter of what its predecessor required.

On the DeepSWE coding benchmark, V4.1-Flash edged out both Opus 5 and GPT-5.6 Sol despite its far smaller active-parameter footprint. DeepSeek has released the model under the permissive MIT license, making it freely available for commercial use and modification.

Why it matters

The headline story here is technical, not experiential: DeepSeek has found a way to shrink the memory overhead that has made running long-context, agentic AI workloads expensive. Key-value cache size is one of the main cost drivers when models hold extended conversation or task history in memory — a quarter of the footprint at comparable or better coding performance is a meaningful efficiency gain, not a marginal one.

For organisations building or deploying AI agents — whether for coding, customer service automation, or internal operations — this lowers the practical cost of running capable models at scale, and an MIT license removes commercial licensing friction. Efficiency gains of this kind tend to widen who can afford to deploy agentic AI in production, rather than just who can prototype with it.

By the numbers

  • 552 billion total parameters in the V4.1-Flash architecture
  • 16 billion parameters actively used per token during inference
  • A quarter of the predecessor's key-value cache memory requirement

The Renascence take

It's tempting to read this purely as a technical benchmark story, but the memory efficiency angle has direct operating-model implications for anyone planning to deploy AI agents at volume rather than in pilot.

Most organisations evaluating AI agents fixate on capability benchmarks and overlook the unit economics of actually running them — memory and inference cost scale with every conversation, every long-context task, every concurrent user. A model that halves or quarters that overhead without sacrificing performance changes the calculus on where agentic AI becomes viable inside a service operation, not just where it's technically impressive. Teams building agent-based customer or employee experiences should be asking their AI vendors and internal teams about cost-per-interaction at scale, not just accuracy scores in a demo.

Sumber-sumber

Taklimat ini ditulis oleh Meja Berita kami, mensintesis laporan daripada saluran di bawah. Ikuti pautan untuk liputan asal.

FAQ

Questions we get on this topic

It's a new multimodal AI model from Chinese developer DeepSeek that uses a mixture-of-experts architecture with 552 billion total parameters but activates only 16 billion per token, designed to make AI agents cheaper to run.

It cuts the memory needed for the key-value cache — one of the main cost drivers for long-context AI workloads — to roughly a quarter of what its predecessor required.

On the DeepSWE coding benchmark, V4.1-Flash outperformed both Opus 5 and GPT-5.6 Sol despite using a far smaller active-parameter footprint.

DeepSeek released the model under the permissive MIT license, making it freely available for commercial use and modification without licensing friction.

Kekal di hadapan CX

Dapatkan isyarat, bukan gangguan.

Kisah-kisah yang membentuk pengalaman pelanggan — serta Jurnal dan Experience Loom — di peti masuk anda.