<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>K3 on mayulu的AI笔记</title><link>https://mayulu.co/tags/k3/</link><description>Recent content in K3 on mayulu的AI笔记</description><generator>Hugo</generator><language>zhcn</language><copyright>Copyright © 2025–2026 mayulu All Rights Reserved</copyright><lastBuildDate>Wed, 29 Jul 2026 10:53:39 +0800</lastBuildDate><atom:link href="https://mayulu.co/tags/k3/index.xml" rel="self" type="application/rss+xml"/><item><title>Sebastian Raschka 对 Kimi K3 架构的观察笔记要点</title><link>https://mayulu.co/posts/rasbt-kimi-k3-arch/</link><pubDate>Wed, 29 Jul 2026 08:53:39 +0800</pubDate><guid>https://mayulu.co/posts/rasbt-kimi-k3-arch/</guid><description>本文全文翻译 Sebastian Raschka 解读的开源超大模型 Kimi K3 架构，该模型由 48B 参数扩容至 2.8T，沿用前代 Kimi Linear 核心设计，新增 LatentMoE 提升效率，搭配多头隐注意力、注意力残差模块优化性能。模型摒弃 RoPE，全域采用 NoPE，原生支持多模态；该设计小幅提升任务效果，仅轻微增加训练与推理开销，是兼顾性能与效率的前沿大模型。</description></item></channel></rss>