Flagos Community

Flagos Community Contact information, map and directions, contact form, opening hours, services, ratings, photos, videos and announcements from Flagos Community, Nonprofit Organization, 150 Chengfu Road, Haidian District, Beijing, Haidian.

FlagOS is a fully open-source AI system software stack for heterogeneous AI chips, allowing AI models to be developed on...
11/03/2026

FlagOS is a fully open-source AI system software stack for heterogeneous AI chips, allowing AI models to be developed once and seamlessly ported to a wide range of AI hardware with minimal effort.

FlagOS comprises four core modules: the operator library FlagGems, the compiler FlagTree, the communication library FlagCX, and the parallel framework FlagScale; and three open-source tools: FlagPerf, FlagRelease, and Triton-Copilot.

Core modules:

FlagGems

FlagGems is a high-performance general-purpose operator library implemented with the Triton programming language and its extended languages. FlagGems is designed to provide a suite of general-purpose operators for large models, accelerating the inference and training of models across multiple backend platforms.

FlagTree

FlagTree is an open-source, unified compiler for multiple AI chips. FlagTree is dedicated to building a compiler and associated tooling platform for diverse AI chips, advancing and expanding the upstream and downstream Triton ecosystem, with the goals of supporting existing adaptation solutions, unifying code repositories, and enabling rapid multi-backend support from a single repository. For upstream model users, FlagTree provides unified compilation support across multiple backends; for downstream chip vendors, FlagTree offers reference implementations for integration into the Triton ecosystem.

FlagScale

FlagScale is a comprehensive toolkit designed to support the entire lifecycle of large models. FlagScale builds on the strengths of several prominent open-source projects, including Megatron-LM and vLLM, to provide a robust, end-to-end solution for managing and scaling large models.

FlagCX

FlagCX is a scalable and adaptive unified communication library for cross-chip environments. FlagCX delivers high-performance point-to-point and collective communication capabilities tailored for multi-chip, multi-platform scenarios. By leveraging the native collective communication capabilities of each platform, FlagCX incorporates technologies such as device-buffer IPC and RDMA to enable highly efficient collective communication in both cross-chip and single-chip scenarios, while also providing adaptive tuning capabilities for communication optimization.

Open-source tools

KernelGen KernelGen is an operator auto-generation tool. KernelGen is designed to construct operator definitions through natural language prompts, retrieve existing similar operator definitions, automatically execute operator accuracy and performance testing, generate accuracy and performance test results, and produce Triton Kernels.

FlagRelease FlagRelease is a platform dedicated to the automatic migration, adaptation and release of large models for multi-architecture AI chips. FlagRelease aims to enable mainstream large models to be migrated, validated, and released on diverse domestic AI hardware with lower cost and higher efficiency through automated, standardized, and intelligent adaptation workflows.

FlagPerf FlagPerf is an integrated AI hardware evaluation engine. FlagPerf aims to establish an industry practice-oriented indicator system and evaluate the actual performance of AI hardware under combinations of software stacks (model + framework + compiler).

Capabilities and advantages of FlagOS
Below are the capabilities and advantages of FlagOS:

Over 20 chip types support and PyTorch and PaddlePaddle support

FlagGEMs supports over 200 core operators.

The unified compiler FlagTree supports chips from 12 vendors.

FlagScale framework enables true cross-framework, cross-chip, and cross-backend unification.

The unified communication library FlagCX supports five communication protocols.

AI-powered automation and development efficiency

KernelGen supports end-to-end operator generation, verification, and optimization across multiple chip backends.

FlagRelease supports continuous deployment of mainstream open-source models on 11 chip types, with automated migration accelerated by FlagOS + AI agents.

Full-stack performance enhancement

Achieves end-to-end inference performance leap, achieves higher speedups than native Triton on key models, approaching the peak performance of CUDA.

FlagScale auto-tuning automatically accelerates both training and inference.

The FlagCX communication library supports pipeline optimization for accelerated performance.

Broader architecture, system, and scenario coverage

Supports more model architectures. Beyond Transformers model, and also supports RWKV and Diffusion-family models.

Supports more system architectures. Compatible with emerging hyperscale node architectures such as Inspur and Hygon.

Supports full-stack embodied intelligence: Covers the entire pipeline—from pre-training and post-training to quantized inference—for “brain” and “cerebellum” (VLA) models, including edge-cloud collaboration and tool retrieval.

Find us more
https://docs.flagos.io/en/latest/overview.html

FlagOS is a fully open-source AI system software stack for heterogeneous AI chips, allowing AI models to be developed once and seamlessly ported to a wide range of AI hardware with minimal effort. Composition modules of FlagOS: The figure below shows the position of FlagOS in the AI ecosystem and...

🚀 FlagOS Global Open Computing Challenge — Season 1 Launch CeremonyWe’re excited to officially launch the FlagOS Global ...
10/03/2026

🚀 FlagOS Global Open Computing Challenge — Season 1 Launch Ceremony

We’re excited to officially launch the FlagOS Global Open Computing Challenge, inviting developers and researchers worldwide to explore high-performance AI computing.

Participants will build and optimize large model operators using the FlagOS platform and compete for a ¥2M prize pool.

📅 Live Stream
Mar 10 | 19:00–20:30 (UTC+8)

🎥 Watch the live launch event:
https://www.youtube.com/live/UigHferD7U0

🚀 FlagOS Global Open Computing Competition – Launch CeremonyJoin us for the official launch of the FlagOS Global Open Computing Competition!This competition...

What is the difference between CANN and FlagOS?CANN focuses on optimizing AI workloads for the Huawei Ascend ecosystem.F...
06/03/2026

What is the difference between CANN and FlagOS?

CANN focuses on optimizing AI workloads for the Huawei Ascend ecosystem.

FlagOS is built for heterogeneous AI computing.
Instead of a single-chip stack, FlagOS supports multiple hardware architectures, including:
- Nvidia GPUs
- Huawei Ascend
- NPUs
- ARM
- RISC-V
It also supports a wide range of multimodal models, such as:
- BAAI Emu
- MiniCPM-V
- Qwen-VL (Qwen2.5 / Qwen3-VL)
- ERNIE 4.5
- LLaVA series
In short:
CANN → single-ecosystem optimization
FlagOS → multi-chip AI infrastructure + multi-model support.

ANN

🚀 FlagOS Lands on Tencent CloudDeploy OpenClaw + LLM on domestic AI chips — and run your own 24/7 AI “digital employee.”...
06/03/2026

🚀 FlagOS Lands on Tencent Cloud
Deploy OpenClaw + LLM on domestic AI chips — and run your own 24/7 AI “digital employee.”
As OpenClaw continues gaining traction, more developers and enterprises are shifting from cloud-only APIs to local AI deployment. Privacy concerns and rising token costs are making on-premise LLM services a real necessity.
Recently, FlagOS partnered with Tencent Cloud HAI to officially launch the Qwen3-4B-hygon-flagos model image on the HAI community platform. Developers can now directly pull and deploy it.
Based on this image, users can quickly run FlagOS + OpenClaw on accelerator cards, using a compact model to power intelligent agents and execute tasks. This enables seamless migration from public cloud APIs to local AI services, while also contributing to the standardization of the domestic AI chip ecosystem.
FlagOS is an open-source AI system software stack developed by the Beijing Academy of Artificial Intelligence. It aims to build a unified, open, and secure full-stack AI platform.
By supporting diverse computing architectures under a unified open-source technology stack, FlagOS enables:
- Develop once, reuse across multiple chips
- Unified deployment across heterogeneous hardware
- Standardized adaptation of domestic AI chips
- Scalable ecosystem growth
The platform supports multiple heterogeneous AI accelerators and helps users rapidly deploy models and AI agents.

🚀 FlagOS + OpenClaw Ultimate Developer Guide Run Qwen3-4B-hygon-flagos locally and unlock private Agent deployment.For y...
27/02/2026

🚀 FlagOS + OpenClaw Ultimate Developer Guide
Run Qwen3-4B-hygon-flagos locally and unlock private Agent deployment.
For years, we’ve defaulted to the cloud for nearly everything. But as Agent frameworks like OpenClaw gain traction, a new demand is emerging:
A 24/7 “digital employee” running on your own infrastructure — fully local, controllable, cost-efficient, and privacy-preserving.
Cloud APIs are powerful — but they also introduce ongoing token costs and potential privacy risks. As Agents continue to scale, token consumption is becoming a structural bottleneck. As a result, more and more enterprises are turning toward localized and private deployment strategies.
That’s where FlagOS comes in.
FlagOS is a fully open-source AI system software stack designed to run models seamlessly across heterogeneous AI hardware. In this tutorial, we demonstrate how to:
✅ Deploy Qwen3-4B-hygon-flagos on Hygon accelerators
✅ Connect it to OpenClaw
✅ Enable tool invocation and system-level control
✅ Integrate with Lark
✅ Validate real-world task ex*****on
🔍 What did we observe?
Small models are no longer just chat components.
Qwen3-4B-hygon-flagos can already handle:
• Instruction understanding
• Tool invocation
• Local file operations
• System coordination
• Controlled enterprise integration
The real bottleneck is shifting — not model size, but system design: permissions, APIs, infrastructure abstraction.
If your goal is to deploy a local Agent core that:
• Runs privately
• Invokes tools reliably
• Integrates with enterprise systems
• Optimizes cost at scale
Then 4B-class models are becoming a practical default choice.
Private AI infrastructure is no longer theoretical. It’s executable.
Less is More. FlagOS is the Key.

🚀 FlagOS Enables Multi-Chip Release of Qwen3.5-397B MoE Unified Multi-Chip Version Now Available for DownloadOn February...
25/02/2026

🚀 FlagOS Enables Multi-Chip Release of Qwen3.5-397B MoE
Unified Multi-Chip Version Now Available for Download
On February 16, 2026 (Chinese New Year’s Eve), Alibaba Cloud released its flagship foundation model — Qwen3.5-397B-A17B
This first large-scale model in the Qwen3.5 series features:
• 397B total parameters
• 17B activated parameters (MoE)
• Native Vision-Language multimodal architecture
It delivers major advancements in:
✔ General intelligence
✔ Code generation
✔ Long-context reasoning
✔ Agent reasoning
✔ Tool use
✔ Multimodal understanding
It is currently one of the most powerful open-source multimodal MoE models available.

---
🚀 Immediate Multi-Chip Adaptation with FlagOS
A 397B MoE model introduces significant system-level challenges:
• Cross-chip backend adaptation
• Multi-node distributed deployment
• Precision alignment
• Performance optimization
Powered by the unified open-source AI system stack
FlagOS
The FlagOS community quickly completed:
✅ Full model adaptation
✅ Precision alignment
✅ Cross-chip migration
Qwen3.5-397B is now simultaneously available on:
• MetaX
• T-Head Zhenwu
• NVIDIA GPUs
---
🔌 Zero Code Changes with vLLM-plugin-FL
Through
vLLM-plugin-FL
Developers can:
• Maintain original vLLM APIs
• Avoid hardware-specific refactoring
• Preserve high-performance inference
• Deploy with zero modification
Validated deployment includes:
• Native BF16
• Dual-node 16-GPU inference
• Verified precision alignment
---
📦 Direct Download – Multi-Chip Ready
Optimized Qwen3.5-FlagOS versions are available on:
• Hugging Face
• ModelScope
All versions are:
✔ Pre-migrated
✔ Pre-validated
✔ Production-ready
✔ No additional adaptation required
Hugging Face: huggingface.co
ModelScope: Qwen3.5-397B-A17B-nvidia-FlagOS
---
Why This Matters
As heterogeneous AI hardware ecosystems expand globally, the real bottleneck is system portability.
FlagOS transforms traditional hardware integration:
From: M × N adaptation complexity
To: M + N unified integration
Large models should not be locked to a single hardware ecosystem.
FlagOS makes them a portable infrastructure.
---
🌍 Join the Ecosystem
FlagOS is organizing the Open Computing Global Challenge
Total Prize Pool: 2,000,000 RMB
Tracks include:
• Operator development
• Inference optimization
• Automatic data annotation
Official site:
https://flagos.io

11/02/2026

🚀 Calling AI infra builders & system engineers.

The FlagOS Open Computing Global Challenge is live with three technical tracks:

Track 1
Low-level operator and kernel implementation with cross-hardware performance optimization.

Track 2
End-to-end inference throughput optimization and system-level performance tuning.

Track 3
Automated data annotation pipeline design for long-context scenarios.

If you’re interested in ML systems, GPU optimization, or scalable AI infrastructure — this is a great opportunity to test your skills.

FlagOS is fully open-source and built for heterogeneous AI accelerators.

Build once. Deploy with minimal adaptation.

🚀 FlagOS Open Computing Global Challenge — Track 3 Now OpenHigh-quality annotated data is the foundation of large langua...
11/02/2026

🚀 FlagOS Open Computing Global Challenge — Track 3 Now Open

High-quality annotated data is the foundation of large language models. Yet for ultra-long text scenarios, traditional human annotation is expensive, slow, and difficult to scale.

Track 3: Automatic Data Annotation for Large Models in Long-Context Scenarios focuses on exploring how large language models can perform accurate, robust, and scalable data annotation under ultra-long context settings, leveraging in-context learning (ICL) and multi-step reasoning.

🔍 What This Track Explores

Automated annotation for ultra-long texts (tens to hundreds of thousands of tokens)

Long-context in-context learning (ICL) frameworks

Example selection, context organization, and interference reduction

Multi-step reasoning for perception, planning, and ex*****on

Collaborative annotation systems approaching human-level quality

All submissions must use FlagScale as the unified framework for LLM loading and ex*****on to ensure fairness and reproducibility.

👥 Who Can Participate

Industry teams

Universities and research institutes

Independent researchers and developers

Teams of up to 3 members are allowed (individual participation welcome).

💻 Compute Resource Support

The organizing committee will provide official compute resources to selected teams.
Applications require a preliminary technical proposal and team background.

📅 Compute resource application deadline: March 11, 2026

🏆 Awards

🥇 First Prize: RMB 30,000
🥈 Second Prize: RMB 20,000 × 3 teams
🥉 Third Prize: RMB 10,000 × 5 teams

All final technical reports and complete source code must be publicly released via the official OpenSeek project page on the FlagOS platform:
🔗 https://flagos.io/RaceDetail?id=296fmsd8&lang=en

🚀 FlagOS Open Computing Global Challenge — Track 2 Now OpenAs large foundation models rapidly scale to tens or even hund...
10/02/2026

🚀 FlagOS Open Computing Global Challenge — Track 2 Now Open

As large foundation models rapidly scale to tens or even hundreds of billions of parameters, achieving high-throughput, low-latency inference under constrained compute resources has become one of the most critical challenges in modern AI systems.

Track 2 of the FlagOS Open Computing Global Challenge invites developers and researchers worldwide to push the limits of large-model inference through end-to-end system-level optimization.

🔍 Track Theme

Large Model Inference Throughput Optimization
Based on the FlagOS Multi-Chip Framework

This track focuses on extreme inference throughput optimization for the Qwen3-4B model under a fixed hardware and software environment, while:

Preserving model accuracy (≤ 2% degradation)

Avoiding noticeable latency regression

Participants are required to optimize inference using the FlagOS inference framework, with opportunities to deeply leverage the FlagGems high-performance operator library.

🔗 Mandatory framework:
https://github.com/flagos-ai/vllm-plugin-FL

🧠 Key Technical Directions

Model, tensor, and multi-chip parallelism

Memory optimization (KV cache compression, dynamic batching)

Compute optimization (operator fusion, custom kernels)

Communication and scheduling optimization

Model compression

Speculative decoding and advanced sampling algorithms

The goal is to fully exploit hardware capabilities and demonstrate state-of-the-art inference system optimization within a unified framework.

🏆 Evaluation & Scoring

70%: Relative throughput improvement over the official baseline

30%: Additional score for each AI chip successfully adapted with verified performance gains

Evaluation is based on official vLLM benchmarking scripts and lm-evaluation-harness, ensuring correctness, fairness, and reproducibility.

👥 Who Should Join

Industry teams

Universities and research institutes

Independent researchers and developers worldwide

Teams may include up to 3 members (solo participation is welcome).

💻 Computing Resource Support

Selected teams may receive unified computing resources from the organizers.
Applicants must submit a preliminary technical proposal and team background.

🗓 Compute resource application deadline: March 11, 2026

⏰ Key Dates

Registration: Jan 9 – May 20

Competition period: Mar 24 – May 20

Results announcement: Early June 2026

💰 Awards

🥇 First Prize: ¥30,000 (1 team)
🥈 Second Prize: ¥20,000 (3 teams)
🥉 Third Prize: ¥10,000 (5 teams)

📌 Submission & Leaderboard
Submit via the FlagOS official platform:
👉 https://flagos.io/RaceDetail?id=296fmr01&lang=en

Top solutions will be merged into the official FlagOS repository via pull requests.

🚀 If you’re working on AI systems, large-model inference, kernel optimization, or multi-chip ex*****on, this challenge is built for real-world system innovation.

Recently, an AMD software executive made a bold statement:“The future of GPU programming is AI coding agents.”Why? Becau...
10/02/2026

Recently, an AMD software executive made a bold statement:
“The future of GPU programming is AI coding agents.”
Why?
Because someone without kernel development experience used an AI agent to port CUDA code to AMD ROCm—in just 30 minutes.
It sounds like magic—but beneath the excitement lies a deeper truth.
Writing code is only part of the problem.
Making that code work well across different AI chips is where things usually break down.
That’s exactly the problem the FlagOS community is tackling.
With KernelGen, AI agents don’t just generate code—they understand operators, verify correctness, benchmark performance, and continuously optimize kernels across hardware platforms. Tasks that once required years of low-level expertise can now be done in hours.
To make this real across ecosystems, KernelGen works with FlagTree, a unified AI compiler that supports nearly 20 AI chip architectures. Together, they reduce hardware migration costs and move us closer to a world where AI software is no longer locked to a single vendor.
AI agents are changing how we write code.
But systems thinking is what changes the rules of the game.

10/02/2026

KernelGen is an operator auto-generation tool that simplifies the entire operator development workflow.
It uses natural language prompts to build operator definitions, retrieves similar existing operators, automatically runs accuracy and performance tests, and generates validated kernels.
🎥 Watch the demo and try KernelGen here:
http://kernelgen.flagos.io/login

Address

150 Chengfu Road, Haidian District, Beijing
Haidian
100084

Alerts

Be the first to know and let us send you an email when Flagos Community posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Shortcuts

Share