复制安装命令
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
[2026/07/22] MNN 3.
用 Codex 或 Claude 安装复制这段 Prompt,粘贴到 Codex、Claude 或其他助手里,让它先审查 Skill 页面再帮你安装。
复制前请先查看来源、License 和安全提示。
来源文件:README.md

MNN is a highly efficient and lightweight deep learning framework. It supports inference and training of deep learning models and has industry-leading performance for inference and training on-device. At present, MNN has been integrated into more than 30 apps of Alibaba Inc, such as Taobao, Tmall, Youku, DingTalk, Xianyu, etc., covering more than 70 usage scenarios such as live broadcast, short video capture, search recommendation, product searching by image, interactive marketing, equity distribution, security risk control. In addition, MNN is also used on embedded devices, such as IoT.
MNN-LLM is a large language model runtime solution developed based on the MNN engine. The mission of this project is to deploy LLM models locally on everyone's platforms(Mobile Phone/PC/IOT). It supports popular large language models such as Qianwen, Baichuan, Zhipu, LLAMA, and others. MNN-LLM User guide
MNN-Diffusion is a stable diffusion model runtime solution developed based on the MNN engine. The mission of this project is to deploy stable diffusion models locally on everyone's platforms. MNN-Diffusion User guide

Inside Alibaba, MNN works as the basic module of the compute container in the Walle System, the first end-to-end, general-purpose, and large-scale production system for device-cloud collaborative machine learning, which has been published in the top system conference OSDI’22. The key design principles of MNN and the extensive benchmark testing results (vs. TensorFlow, TensorFlow Lite, PyTorch, PyTorch Mobile, TVM) can be found in the OSDI paper. The scripts and instructions for benchmark testing are put in the path “/benchmark”. If MNN or the design of Walle helps your research or production use, please cite our OSDI paper as follows:
@inproceedings {proc:osdi22:walle,
author = {Chengfei Lv and Chaoyue Niu and Renjie Gu and Xiaotang Jiang and Zhaode Wang and Bin Liu and Ziqi Wu and Qiulin Yao and Congyu Huang and Panos Huang and Tao Huang and Hui Shu and Jinde Song and Bin Zou and Peng Lan and Guohuan Xu and Fei Wu and Shaojie Tang and Fan Wu and Guihai Chen},
title = {Walle: An {End-to-End}, {General-Purpose}, and {Large-Scale} Production System for {Device-Cloud} Collaborative Machine Learning},
booktitle = {16th USENIX Symposium on Operating Systems Design and Implementation (OSDI 22)},
year = {2022},
isbn = {978-1-939133-28-1},
address = {Carlsbad, CA},
pages = {249--265},
url = {https://www.usenix.org/conference/osdi22/presentation/lv},
publisher = {USENIX Association},
month = jul,
}
MNN's docs are in place in Read the docs.
You can also read docs/README to build docs's html.
MNN Workbench could be downloaded from MNN's homepage, which provides pretrained models, visualized training tools, and one-click deployment of models to devices.
Tensorflow, Caffe, ONNX,Torchscripts and supports common neural networks such as CNN, RNN, GAN, Transformer.Tensorflow OPs, 52 Caffe OPs, 163 Torchscripts OPs, 158 ONNX OPs.The Architecture / Precision MNN supported is shown below:
| Architecture / Precision | Normal | FP16 | BF16 | Int8 | |
|---|---|---|---|---|---|
| CPU | Native | B | C | B | B |
| x86/x64-SSE4.1 | A | C | C | A | |
| x86/x64-AVX2 | S | C | C | A | |
| x86/x64-AVX512 | S | C | C | S | |
| ARMv7a | S | S (ARMv8.2) | S | S | |
| ARMv8 | S | S (ARMv8.2) | S(ARMv8.6) | S | |
| GPU | OpenCL | A | S | C | S |
| Vulkan | A | A | C | A | |
| Metal | A | S | C | S | |
| CUDA | A | S | C | A | |
| NPU | CoreML | A | C | C | C |
| HIAI | A | C | C | C | |
| NNAPI | B | B | C | B | |
| QNN | C | B | C | C |
Base on MNN (Tensor compute engine), we provided a series of tools for inference, train and general computation.
The group discussions are predominantly Chinese. But we welcome and will help English speakers.
Dingtalk discussion groups:
Group #4 (Available): 160170007549
Group #3 (Full)
Group #2 (Full): 23350225
Group #1 (Full): 23329087
The preliminary version of MNN, as mobile inference engine and with the focus on manual optimization, has also been published in MLSys 2020. Please cite the paper, if MNN previously helped your research:
@inproceedings{alibaba2020mnn,
author = {Jiang, Xiaotang and Wang, Huan and Chen, Yiliu and Wu, Ziqi and Wang, Lichuan and Zou, Bin and Yang, Yafeng and Cui, Zongyang and Cai, Yu and Yu, Tianhang and Lv, Chengfei and Wu, Zhihua},
title = {MNN: A Universal and Efficient Inference Engine},
booktitle = {MLSys},
year = {2020}
}
Apache 2.0
MNN participants: Taobao Technology Department, Search Engineering Team, DAMO Team, Youku and other Alibaba Group employees.
MNN refers to the following projects:
name: retrospective
description: 任务完成后的反思与经验沉淀。仅在任务非平凡且出现调试、失败修复、反复试错、明显误判或可复用教训时使用;简单执行、查询、常规 CI 通过等无新增经验的任务可跳过。触发条件:非平凡任务完成后,且存在可沉淀经验:debug、失败修复、反复试错、明显误判、流程缺口,或用户显式要求。
跳过条件:简单执行、查询、格式调整、常规 CI 通过等没有新增教训的任务,无需执行此 skill。
服务对象:所有 skill。本 skill 的产出是对其他 skill 文件的更新,不产生代码。
最大的风险是"觉得差不多了"的幻觉——结果不达标时,找外部原因搪塞然后降低要求宣布完成。
常见的自我欺骗模式:
正确做法:
每个子任务开始前明确三件事:
| 教训类型 | 更新位置 |
|---|---|
| 某个 skill 流程中的陷阱 | 该 skill 的 common-pitfalls 或步骤文件 |
| 通用工作方法论 | 该 skill 的 SKILL.md 顶层注意事项 |
| 跨 skill 通用经验 | memory 系统 |
评论 (0)
暂无评论,成为第一个评论者吧!