# Data Leads Future > One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Every article saves you 40+ hours of work. Public Ghost content for AI and LLM tooling. This file includes a bounded export of public pages first, then recent public posts. Append `.md` to any post or page URL to get the content in Markdown (for example, `/example-post.md`). ## Pages ### About this site URL: https://www.dataleadsfuture.com/about/ Last updated: 2026-07-09T08:21:16.000Z ## About me ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/avatar-1.png) **Who am I** Hi, I'm [Peng Qian](https://www.linkedin.com/in/qtalen/?ref=dataleadsfuture.com). I was a senior data scientist at Alibaba. Now I'm the chief architect of the big data department at China Merchants Securities, one of China's largest investment banks. I lead research on the integration of data and AI. **What I do:** I don't want data or AI to be toys for a small group of "PhDs." I want you to use AI and data science to make your work easier. So I'm happy to share my hands-on experience in data science and AI, and grow together with you. **Why am I here:** There's a saying: the best investment is investing in yourself. Through learning and practice, I believe we can all grow in the tech field. The best way to learn is to share. I think blogs are the best way to share: easy to search, and full of text and charts. So I created Data Leads Future to share how I use data science and AI to solve real problems at work. I'm also eager to learn from you and grow together. --- ## About content Data Leads Future is dedicated to sharing data science stories from real projects, each containing unique and practical project experiences. I solemnly promise: Every article I write comes from my own fingers typing word by word on the keyboard. Even if some of the views in my articles turn out to be wrong, they still represent my genuine voice. **I don't use any AI tools to generate my article content.** However, since I come from a non-native English-speaking country, to make my articles easier to understand, I do use AI to help me create article cover images (DALL-E-3, gpt-image-2), and I also use AI for translation and grammar proofreading (claude-sonnet-4.6, Grammarly). I hope you don't mind. --- ## About Data Leads Future Data Leads Future is an independent publication launched in August 2023 by me. Which aims to share knowledge on how to solve real-world problems using AI and data science, for everyone from beginners to experts. If you subscribe today, you'll get full access to the website as well as email newsletters about new content when it's available. Your subscription makes this site possible and allows Data Leads Future to continue to exist. Thank you! --- ## Access all areas By signing up, you'll get access to the full archive of everything that's been published before and everything that's still to come. Your very own private library. ## Fresh content, delivered Stay up to date with new content sent straight to your inbox! No more worrying about whether you missed something because of a pesky algorithm or news feed. ## Meet people like you Join a community of other subscribers who share the same interests. ## Contact You can find me on: [LinkedIn](https://www.linkedin.com/in/qtalen/?ref=dataleadsfuture.com). If you have any questions about Data Leads Future, you can contact me at [dataleadsfuture@gmail.com](mailto:dataleadsfuture@gmail.com). My Zhihu column: [数据引领未来](https://www.zhihu.com/column/c%5F1685309828812107776?ref=dataleadsfuture.com) My WeChat official account is: Data Leads Future ### Privacy Policy URL: https://www.dataleadsfuture.com/privacy-policy/ Last updated: 2023-08-27T06:07:29.000Z **PRIVACY POLICY** **Last updated August 27, 2023** This privacy notice for Data Leads Future ("**we**," "**us**," or "**our**"), describes how and why we might collect, store, use, and/or share ("**process**") your information when you use our services ("**Services**"), such as when you: - Visit our website at [https://www.dataleadsfuture.com](https://www.dataleadsfuture.com/), or any website of ours that links to this privacy notice - Engage with us in other related ways, including any sales, marketing, or events **Questions or concerns?** Reading this privacy notice will help you understand your privacy rights and choices. If you do not agree with our policies and practices, please do not use our Services. If you still have any questions or concerns, please contact us at dataleadsfuture@gmail.com. **SUMMARY OF KEY POINTS** ***This summary provides key points from our privacy notice, but you can find out more details about any of these topics by clicking the link following each key point or by using our*** [***table of contents***](#toc) ***below to find the section you are looking for.*** **What personal information do we process?** When you visit, use, or navigate our Services, we may process personal information depending on how you interact with us and the Services, the choices you make, and the products and features you use. Learn more about [personal information you disclose to us](#personalinfo). **Do we process any sensitive personal information?** We do not process sensitive personal information. **Do we receive any information from third parties?** We may receive information from public databases, marketing partners, social media platforms, and other outside sources. Learn more about [information collected from other sources](#othersources). **How do we process your information?** We process your information to provide, improve, and administer our Services, communicate with you, for security and fraud prevention, and to comply with law. We may also process your information for other purposes with your consent. We process your information only when we have a valid legal reason to do so. Learn more about [how we process your information](#infouse). **In what situations and with which types of parties do we share personal information?** We may share information in specific situations and with specific categories of third parties. Learn more about [when and with whom we share your personal information](#whoshare). **How do we keep your information safe?** We have organizational and technical processes and procedures in place to protect your personal information. However, no electronic transmission over the internet or information storage technology can be guaranteed to be 100% secure, so we cannot promise or guarantee that hackers, cybercriminals, or other unauthorized third parties will not be able to defeat our security and improperly collect, access, steal, or modify your information. Learn more about [how we keep your information safe](#infosafe). **What are your rights?** Depending on where you are located geographically, the applicable privacy law may mean you have certain rights regarding your personal information. Learn more about [your privacy rights](#privacyrights). **How do you exercise your rights?** The easiest way to exercise your rights is by submitting a [data subject access request](https://app.termly.io/notify/489d8eac-3a1d-4368-bc83-aa3cb2203283?ref=dataleadsfuture.com), or by contacting us. We will consider and act upon any request in accordance with applicable data protection laws. Want to learn more about what we do with any information we collect? [Review the privacy notice in full](#toc). **TABLE OF CONTENTS** [1\. WHAT INFORMATION DO WE COLLECT?](#infocollect) [2\. HOW DO WE PROCESS YOUR INFORMATION?](#infouse) [3\. WHAT LEGAL BASES DO WE RELY ON TO PROCESS YOUR PERSONAL INFORMATION?](#legalbases) [4\. WHEN AND WITH WHOM DO WE SHARE YOUR PERSONAL INFORMATION?](#whoshare) [5\. WHAT IS OUR STANCE ON THIRD-PARTY WEBSITES?](#3pwebsites) [6\. DO WE USE COOKIES AND OTHER TRACKING TECHNOLOGIES?](#cookies) [7\. IS YOUR INFORMATION TRANSFERRED INTERNATIONALLY?](#intltransfers) [8\. HOW LONG DO WE KEEP YOUR INFORMATION?](#inforetain) [9\. HOW DO WE KEEP YOUR INFORMATION SAFE?](#infosafe) [10\. DO WE COLLECT INFORMATION FROM MINORS?](#infominors) [11\. WHAT ARE YOUR PRIVACY RIGHTS?](#privacyrights) [12\. CONTROLS FOR DO-NOT-TRACK FEATURES](#DNT) [13\. DO CALIFORNIA RESIDENTS HAVE SPECIFIC PRIVACY RIGHTS?](#caresidents) [14\. DO VIRGINIA RESIDENTS HAVE SPECIFIC PRIVACY RIGHTS?](#virginia) [15\. DO WE MAKE UPDATES TO THIS NOTICE?](#policyupdates) [16\. HOW CAN YOU CONTACT US ABOUT THIS NOTICE?](#contact) [17\. HOW CAN YOU REVIEW, UPDATE, OR DELETE THE DATA WE COLLECT FROM YOU?](#request) **1\. WHAT INFORMATION DO WE COLLECT?** **Personal information you disclose to us** ***In Short:*** *We collect personal information that you provide to us.* We collect personal information that you voluntarily provide to us when you register on the Services, express an interest in obtaining information about us or our products and Services, when you participate in activities on the Services, or otherwise when you contact us. **Personal Information Provided by You.** The personal information that we collect depends on the context of your interactions with us and the Services, the choices you make, and the products and features you use. The personal information we collect may include the following: - email addresses - job titles **Sensitive Information.** We do not process sensitive information. **Payment Data.** We may collect data necessary to process your payment if you make purchases, such as your payment instrument number, and the security code associated with your payment instrument. All payment data is stored by Stripe. You may find their privacy notice link(s) here: [https://stripe.com/zh-cn-hk/privacy](https://stripe.com/zh-cn-hk/privacy?ref=dataleadsfuture.com). All personal information that you provide to us must be true, complete, and accurate, and you must notify us of any changes to such personal information. **Information automatically collected** ***In Short:*** *Some information — such as your Internet Protocol (IP) address and/or browser and device characteristics — is collected automatically when you visit our Services.* We automatically collect certain information when you visit, use, or navigate the Services. This information does not reveal your specific identity (like your name or contact information) but may include device and usage information, such as your IP address, browser and device characteristics, operating system, language preferences, referring URLs, device name, country, location, information about how and when you use our Services, and other technical information. This information is primarily needed to maintain the security and operation of our Services, and for our internal analytics and reporting purposes. Like many businesses, we also collect information through cookies and similar technologies. You can find out more about this in our Cookie Notice: [http://www.dataleadsfuture.com/cookies](https://www.dataleadsfuture.com/cookies). The information we collect includes: - *Log and Usage Data.* Log and usage data is service-related, diagnostic, usage, and performance information our servers automatically collect when you access or use our Services and which we record in log files. Depending on how you interact with us, this log data may include your IP address, device information, browser type, and settings and information about your activity in the Services (such as the date/time stamps associated with your usage, pages and files viewed, searches, and other actions you take such as which features you use), device event information (such as system activity, error reports (sometimes called "crash dumps" ), and hardware settings). - *Device Data.* We collect device data such as information about your computer, phone, tablet, or other device you use to access the Services. Depending on the device used, this device data may include information such as your IP address (or proxy server), device and application identification numbers, location, browser type, hardware model, Internet service provider and/or mobile carrier, operating system, and system configuration information. - *Location Data.* We collect location data such as information about your device's location, which can be either precise or imprecise. How much information we collect depends on the type and settings of the device you use to access the Services. For example, we may use GPS and other technologies to collect geolocation data that tells us your current location (based on your IP address). You can opt out of allowing us to collect this information either by refusing access to the information or by disabling your Location setting on your device. However, if you choose to opt out, you may not be able to use certain aspects of the Services. **Information collected from other sources** ***In Short:*** *We may collect limited data from public databases, marketing partners, and other outside sources.* In order to enhance our ability to provide relevant marketing, offers, and services to you and update our records, we may obtain information about you from other sources, such as public databases, joint marketing partners, affiliate programs, data providers, and from other third parties. This information includes mailing addresses, job titles, email addresses, phone numbers, intent data (or user behavior data), Internet Protocol (IP) addresses, social media profiles, social media URLs, and custom profiles, for purposes of targeted advertising and event promotion. **2\. HOW DO WE PROCESS YOUR INFORMATION?** ***In Short:*** *We process your information to provide, improve, and administer our Services, communicate with you, for security and fraud prevention, and to comply with law. We may also process your information for other purposes with your consent.* **We process your personal information for a variety of reasons, depending on how you interact with our Services, including:** - **To facilitate account creation and authentication and otherwise manage user accounts.** We may process your information so you can create and log in to your account, as well as keep your account in working order. - **To deliver and facilitate delivery of services to the user.** We may process your information to provide you with the requested service. - **To respond to user inquiries/offer support to users.** We may process your information to respond to your inquiries and solve any potential issues you might have with the requested service. - **To send administrative information to you.** We may process your information to send you details about our products and services, changes to our terms and policies, and other similar information. - **To fulfill and manage your orders.** We may process your information to fulfill and manage your orders, payments, returns, and exchanges made through the Services. - **To request feedback.** We may process your information when necessary to request feedback and to contact you about your use of our Services. - **To protect our Services.** We may process your information as part of our efforts to keep our Services safe and secure, including fraud monitoring and prevention. - **To save or protect an individual's vital interest.** We may process your information when necessary to save or protect an individual’s vital interest, such as to prevent harm. **3\. WHAT LEGAL BASES DO WE RELY ON TO PROCESS YOUR INFORMATION?** **In Short:* We only process your personal information when we believe it is necessary and we have a valid legal reason (i.e., legal basis) to do so under applicable law, like with your consent, to comply with laws, to provide you with services to enter into or fulfill our contractual obligations, to protect your rights, or to fulfill our legitimate business interests.* **If you are located in the EU or UK, this section applies to you.** The General Data Protection Regulation (GDPR) and UK GDPR require us to explain the valid legal bases we rely on in order to process your personal information. As such, we may rely on the following legal bases to process your personal information: - **Consent.** We may process your information if you have given us permission (i.e., consent) to use your personal information for a specific purpose. You can withdraw your consent at any time. Learn more about [withdrawing your consent](#withdrawconsent). - **Performance of a Contract.** We may process your personal information when we believe it is necessary to fulfill our contractual obligations to you, including providing our Services or at your request prior to entering into a contract with you. - **Legitimate Interests.** We may process your information when we believe it is reasonably necessary to achieve our legitimate business interests and those interests do not outweigh your interests and fundamental rights and freedoms. For example, we may process your personal information for some of the purposes described in order to: - Diagnose problems and/or prevent fraudulent activities - Understand how our users use our products and services so we can improve user experience - **Legal Obligations.** We may process your information where we believe it is necessary for compliance with our legal obligations, such as to cooperate with a law enforcement body or regulatory agency, exercise or defend our legal rights, or disclose your information as evidence in litigation in which we are involved. - **Vital Interests.** We may process your information where we believe it is necessary to protect your vital interests or the vital interests of a third party, such as situations involving potential threats to the safety of any person. ***If you are located in Canada, this section applies to you.*** We may process your information if you have given us specific permission (i.e., express consent) to use your personal information for a specific purpose, or in situations where your permission can be inferred (i.e., implied consent). You can [withdraw your consent](#withdrawconsent) at any time. In some exceptional cases, we may be legally permitted under applicable law to process your information without your consent, including, for example: - If collection is clearly in the interests of an individual and consent cannot be obtained in a timely way - For investigations and fraud detection and prevention - For business transactions provided certain conditions are met - If it is contained in a witness statement and the collection is necessary to assess, process, or settle an insurance claim - For identifying injured, ill, or deceased persons and communicating with next of kin - If we have reasonable grounds to believe an individual has been, is, or may be victim of financial abuse - If it is reasonable to expect collection and use with consent would compromise the availability or the accuracy of the information and the collection is reasonable for purposes related to investigating a breach of an agreement or a contravention of the laws of Canada or a province - If disclosure is required to comply with a subpoena, warrant, court order, or rules of the court relating to the production of records - If it was produced by an individual in the course of their employment, business, or profession and the collection is consistent with the purposes for which the information was produced - If the collection is solely for journalistic, artistic, or literary purposes - If the information is publicly available and is specified by the regulations **4\. WHEN AND WITH WHOM DO WE SHARE YOUR PERSONAL INFORMATION?** ***In Short:*** *We may share information in specific situations described in this section and/or with the following categories of third parties.* **Vendors, Consultants, and Other Third-Party Service Providers.** We may share your data with third-party vendors, service providers, contractors, or agents ("**third parties**") who perform services for us or on our behalf and require access to such information to do that work. We have contracts in place with our third parties, which are designed to help safeguard your personal information. This means that they cannot do anything with your personal information unless we have instructed them to do it. They will also not share your personal information with any organization apart from us. They also commit to protect the data they hold on our behalf and to retain it for the period we instruct. The categories of third parties we may share personal information with are as follows: - Ad Networks - Affiliate Marketing Programs - Data Analytics Services - Cloud Computing Services - Communication & Collaboration Tools - Data Storage Service Providers - Order Fulfillment Service Providers - Payment Processors - Performance Monitoring Tools - Product Engineering & Design Tools - Retargeting Platforms - Testing Tools - User Account Registration & Authentication Services - Website Hosting Service Providers - Sales & Marketing Tools We also may need to share your personal information in the following situations: - **Business Transfers.** We may share or transfer your information in connection with, or during negotiations of, any merger, sale of company assets, financing, or acquisition of all or a portion of our business to another company. - **Affiliates.** We may share your information with our affiliates, in which case we will require those affiliates to honor this privacy notice. Affiliates include our parent company and any subsidiaries, joint venture partners, or other companies that we control or that are under common control with us. - **Business Partners.** We may share your information with our business partners to offer you certain products, services, or promotions. **5\. WHAT IS OUR STANCE ON THIRD-PARTY WEBSITES?** ***In Short:*** *We are not responsible for the safety of any information that you share with third parties that we may link to or who advertise on our Services, but are not affiliated with, our Services.* The Services may link to third-party websites, online services, or mobile applications and/or contain advertisements from third parties that are not affiliated with us and which may link to other websites, services, or applications. Accordingly, we do not make any guarantee regarding any such third parties, and we will not be liable for any loss or damage caused by the use of such third-party websites, services, or applications. The inclusion of a link towards a third-party website, service, or application does not imply an endorsement by us. We cannot guarantee the safety and privacy of data you provide to any third parties. Any data collected by third parties is not covered by this privacy notice. We are not responsible for the content or privacy and security practices and policies of any third parties, including other websites, services, or applications that may be linked to or from the Services. You should review the policies of such third parties and contact them directly to respond to your questions. **6\. DO WE USE COOKIES AND OTHER TRACKING TECHNOLOGIES?** ***In Short:*** *We may use cookies and other tracking technologies to collect and store your information.* We may use cookies and similar tracking technologies (like web beacons and pixels) to access or store information. Specific information about how we use such technologies and how you can refuse certain cookies is set out in our Cookie Notice: [http://www.dataleadsfuture.com/cookies](https://www.dataleadsfuture.com/cookies). **7\. IS YOUR INFORMATION TRANSFERRED INTERNATIONALLY?** ***In Short:*** *We may transfer, store, and process your information in countries other than your own.* Our servers are located in the United States. If you are accessing our Services from outside the United States, please be aware that your information may be transferred to, stored, and processed by us in our facilities and by those third parties with whom we may share your personal information (see "[WHEN AND WITH WHOM DO WE SHARE YOUR PERSONAL INFORMATION?](#whoshare)" above), in the United States, and other countries. If you are a resident in the European Economic Area (EEA) or United Kingdom (UK), then these countries may not necessarily have data protection laws or other similar laws as comprehensive as those in your country. However, we will take all necessary measures to protect your personal information in accordance with this privacy notice and applicable law. European Commission's Standard Contractual Clauses: We have implemented measures to protect your personal information, including by using the European Commission's Standard Contractual Clauses for transfers of personal information between our group companies and between us and our third-party providers. These clauses require all recipients to protect all personal information that they process originating from the EEA or UK in accordance with European data protection laws and regulations. Our Standard Contractual Clauses can be provided upon request. We have implemented similar appropriate safeguards with our third-party service providers and partners and further details can be provided upon request. **8\. HOW LONG DO WE KEEP YOUR INFORMATION?** ***In Short:*** *We keep your information for as long as necessary to fulfill the purposes outlined in this privacy notice unless otherwise required by law.* We will only keep your personal information for as long as it is necessary for the purposes set out in this privacy notice, unless a longer retention period is required or permitted by law (such as tax, accounting, or other legal requirements). No purpose in this notice will require us keeping your personal information for longer than three (3) months past the termination of the user's account. When we have no ongoing legitimate business need to process your personal information, we will either delete or anonymize such information, or, if this is not possible (for example, because your personal information has been stored in backup archives), then we will securely store your personal information and isolate it from any further processing until deletion is possible. **9\. HOW DO WE KEEP YOUR INFORMATION SAFE?** ***In Short:*** *We aim to protect your personal information through a system of organizational and technical security measures.* We have implemented appropriate and reasonable technical and organizational security measures designed to protect the security of any personal information we process. However, despite our safeguards and efforts to secure your information, no electronic transmission over the Internet or information storage technology can be guaranteed to be 100% secure, so we cannot promise or guarantee that hackers, cybercriminals, or other unauthorized third parties will not be able to defeat our security and improperly collect, access, steal, or modify your information. Although we will do our best to protect your personal information, transmission of personal information to and from our Services is at your own risk. You should only access the Services within a secure environment. **10\. DO WE COLLECT INFORMATION FROM MINORS?** ***In Short:*** *We do not knowingly collect data from or market to children under 18 years of age.* We do not knowingly solicit data from or market to children under 18 years of age. By using the Services, you represent that you are at least 18 or that you are the parent or guardian of such a minor and consent to such minor dependent’s use of the Services. If we learn that personal information from users less than 18 years of age has been collected, we will deactivate the account and take reasonable measures to promptly delete such data from our records. If you become aware of any data we may have collected from children under age 18, please contact us at \_\_\_\_\_\_\_\_\_\_. **11\. WHAT ARE YOUR PRIVACY RIGHTS?** ***In Short:*** *In some regions, such as the European Economic Area (EEA), United Kingdom (UK), and Canada, you have rights that allow you greater access to and control over your personal information. You may review, change, or terminate your account at any time.* In some regions (like the EEA, UK, and Canada), you have certain rights under applicable data protection laws. These may include the right (i) to request access and obtain a copy of your personal information, (ii) to request rectification or erasure; (iii) to restrict the processing of your personal information; and (iv) if applicable, to data portability. In certain circumstances, you may also have the right to object to the processing of your personal information. You can make such a request by contacting us by using the contact details provided in the section "[HOW CAN YOU CONTACT US ABOUT THIS NOTICE?](#contact)" below. We will consider and act upon any request in accordance with applicable data protection laws. If you are located in the EEA or UK and you believe we are unlawfully processing your personal information, you also have the right to complain to your [Member State data protection authority](https://ec.europa.eu/justice/data-protection/bodies/authorities/index%5Fen.htm?ref=dataleadsfuture.com) or [UK data protection authority](https://ico.org.uk/make-a-complaint/data-protection-complaints/data-protection-complaints/?ref=dataleadsfuture.com). If you are located in Switzerland, you may contact the [Federal Data Protection and Information Commissioner](https://www.edoeb.admin.ch/edoeb/en/home.html?ref=dataleadsfuture.com). **Withdrawing your consent:** If we are relying on your consent to process your personal information, which may be express and/or implied consent depending on the applicable law, you have the right to withdraw your consent at any time. You can withdraw your consent at any time by contacting us by using the contact details provided in the section "[HOW CAN YOU CONTACT US ABOUT THIS NOTICE?](#contact)" below or updating your preferences. However, please note that this will not affect the lawfulness of the processing before its withdrawal nor, when applicable law allows, will it affect the processing of your personal information conducted in reliance on lawful processing grounds other than consent. **Opting out of marketing and promotional communications:**You can unsubscribe from our marketing and promotional communications at any time by clicking on the unsubscribe link in the emails that we send, replying "STOP" or "UNSUBSCRIBE" to the SMS messages that we send, or by contacting us using the details provided in the section "[HOW CAN YOU CONTACT US ABOUT THIS NOTICE?](#contact)" below. You will then be removed from the marketing lists. However, we may still communicate with you — for example, to send you service-related messages that are necessary for the administration and use of your account, to respond to service requests, or for other non-marketing purposes. **Account Information** If you would at any time like to review or change the information in your account or terminate your account, you can: - Log in to your account settings and update your user account. - Contact us using the contact information provided. Upon your request to terminate your account, we will deactivate or delete your account and information from our active databases. However, we may retain some information in our files to prevent fraud, troubleshoot problems, assist with any investigations, enforce our legal terms and/or comply with applicable legal requirements. **Cookies and similar technologies:** Most Web browsers are set to accept cookies by default. If you prefer, you can usually choose to set your browser to remove cookies and to reject cookies. If you choose to remove cookies or reject cookies, this could affect certain features or services of our Services. You may also [opt out of interest-based advertising by advertisers](http://www.aboutads.info/choices/?ref=dataleadsfuture.com) on our Services. For further information, please see our Cookie Notice: [http://www.dataleadsfuture.com/cookies](https://www.dataleadsfuture.com/cookies). If you have questions or comments about your privacy rights, you may email us at dataleadsfuture@gmail.com. **12\. CONTROLS FOR DO-NOT-TRACK FEATURES** Most web browsers and some mobile operating systems and mobile applications include a Do-Not-Track ("DNT") feature or setting you can activate to signal your privacy preference not to have data about your online browsing activities monitored and collected. At this stage no uniform technology standard for recognizing and implementing DNT signals has been finalized. As such, we do not currently respond to DNT browser signals or any other mechanism that automatically communicates your choice not to be tracked online. If a standard for online tracking is adopted that we must follow in the future, we will inform you about that practice in a revised version of this privacy notice. **13\. DO CALIFORNIA RESIDENTS HAVE SPECIFIC PRIVACY RIGHTS?** ***In Short:*** *Yes, if you are a resident of California, you are granted specific rights regarding access to your personal information.* California Civil Code Section 1798.83, also known as the "Shine The Light" law, permits our users who are California residents to request and obtain from us, once a year and free of charge, information about categories of personal information (if any) we disclosed to third parties for direct marketing purposes and the names and addresses of all third parties with which we shared personal information in the immediately preceding calendar year. If you are a California resident and would like to make such a request, please submit your request in writing to us using the contact information provided below. If you are under 18 years of age, reside in California, and have a registered account with Services, you have the right to request removal of unwanted data that you publicly post on the Services. To request removal of such data, please contact us using the contact information provided below and include the email address associated with your account and a statement that you reside in California. We will make sure the data is not publicly displayed on the Services, but please be aware that the data may not be completely or comprehensively removed from all our systems (e.g., backups, etc.). **CCPA Privacy Notice** The California Code of Regulations defines a "resident" as: (1) every individual who is in the State of California for other than a temporary or transitory purpose and (2) every individual who is domiciled in the State of California who is outside the State of California for a temporary or transitory purpose All other individuals are defined as "non-residents." If this definition of "resident" applies to you, we must adhere to certain rights and obligations regarding your personal information. **What categories of personal information do we collect?** We have collected the following categories of personal information in the past twelve (12) months: **Category** **Examples** **Collected** A. Identifiers Contact details, such as real name, alias, postal address, telephone or mobile contact number, unique personal identifier, online identifier, Internet Protocol address, email address, and account name YES B. Personal information categories listed in the California Customer Records statute Name, contact information, education, employment, employment history, and financial information NO C. Protected classification characteristics under California or federal law Gender and date of birth NO D. Commercial information Transaction information, purchase history, financial details, and payment information NO E. Biometric information Fingerprints and voiceprints NO F. Internet or other similar network activity Browsing history, search history, online behavior, interest data, and interactions with our and other websites, applications, systems, and advertisements YES G. Geolocation data Device location YES H. Audio, electronic, visual, thermal, olfactory, or similar information Images and audio, video or call recordings created in connection with our business activities NO I. Professional or employment-related information Business contact details in order to provide you our Services at a business level or job title, work history, and professional qualifications if you apply for a job with us NO J. Education Information Student records and directory information NO K. Inferences drawn from other personal information Inferences drawn from any of the collected personal information listed above to create a profile or summary about, for example, an individual’s preferences and characteristics YES L. Sensitive Personal Information NO We will use and retain the collected personal information as needed to provide the Services or for: - Category A - As long as the user has an account with us - Category F - 1 year - Category G - 1 year - Category K - 1 year We may also collect other personal information outside of these categories through instances where you interact with us in person, online, or by phone or mail in the context of: - Receiving help through our customer support channels; - Participation in customer surveys or contests; and - Facilitation in the delivery of our Services and to respond to your inquiries. **How do we use and share your personal information?** We collect and share your personal information through: - Targeting cookies/Marketing cookies More information about our data collection and sharing practices can be found in this privacy notice and our Cookie Notice: [http://www.dataleadsfuture.com/cookies](https://www.dataleadsfuture.com/cookies). You may contact us by email at dataleadsfuture@gmail.com, or by referring to the contact details at the bottom of this document. If you are using an authorized agent to exercise your right to opt out we may deny a request if the authorized agent does not submit proof that they have been validly authorized to act on your behalf. **Will your information be shared with anyone else?** We may disclose your personal information with our service providers pursuant to a written contract between us and each service provider. Each service provider is a for-profit entity that processes the information on our behalf, following the same strict privacy protection obligations mandated by the CCPA. We may use your personal information for our own business purposes, such as for undertaking internal research for technological development and demonstration. This is not considered to be "selling" of your personal information. We have not sold or shared any personal information to third parties for a business or commercial purpose in the preceding twelve (12) months. We have disclosed the following categories of personal information to third parties for a business or commercial purpose in the preceding twelve (12) months: - Category A. Identifiers, such as contact details like your real name, alias, postal address, telephone or mobile contact number, unique personal identifier, online identifier, Internet Protocol address, email address, and account name. Category F. Internet or other electronic network activity information, such as browsing history, search history, online behavior, interest data, and interactions with our and other websites, applications, systems, and advertisements. - Category G. Geolocation data, such as device location. - Category K. Inferences drawn from any of the personal information listed above to create a profile or summary about, for example, an individual's preferences and characteristics. The categories of third parties to whom we disclosed personal information for a business or commercial purpose can be found under "[WHEN AND WITH WHOM DO WE SHARE YOUR PERSONAL INFORMATION?](#whoshare)". **Your rights with respect to your personal data** Right to request deletion of the data — Request to delete You can ask for the deletion of your personal information. If you ask us to delete your personal information, we will respect your request and delete your personal information, subject to certain exceptions provided by law, such as (but not limited to) the exercise by another consumer of his or her right to free speech, our compliance requirements resulting from a legal obligation, or any processing that may be required to protect against illegal activities. Right to be informed — Request to know Depending on the circumstances, you have a right to know: - whether we collect and use your personal information; - the categories of personal information that we collect; - the purposes for which the collected personal information is used; - whether we sell or share personal information to third parties; - the categories of personal information that we sold, shared, or disclosed for a business purpose; - the categories of third parties to whom the personal information was sold, shared, or disclosed for a business purpose; - the business or commercial purpose for collecting, selling, or sharing personal information; and - the specific pieces of personal information we collected about you. In accordance with applicable law, we are not obligated to provide or delete consumer information that is de-identified in response to a consumer request or to re-identify individual data to verify a consumer request. Right to Non-Discrimination for the Exercise of a Consumer’s Privacy Rights We will not discriminate against you if you exercise your privacy rights. Right to Limit Use and Disclosure of Sensitive Personal Information We do not process consumer's sensitive personal information. Verification process Upon receiving your request, we will need to verify your identity to determine you are the same person about whom we have the information in our system. These verification efforts require us to ask you to provide information so that we can match it with information you have previously provided us. For instance, depending on the type of request you submit, we may ask you to provide certain information so that we can match the information you provide with the information we already have on file, or we may contact you through a communication method (e.g., phone or email) that you have previously provided to us. We may also use other verification methods as the circumstances dictate. We will only use personal information provided in your request to verify your identity or authority to make the request. To the extent possible, we will avoid requesting additional information from you for the purposes of verification. However, if we cannot verify your identity from the information already maintained by us, we may request that you provide additional information for the purposes of verifying your identity and for security or fraud-prevention purposes. We will delete such additionally provided information as soon as we finish verifying you. Other privacy rights - You may object to the processing of your personal information. - You may request correction of your personal data if it is incorrect or no longer relevant, or ask to restrict the processing of the information. - You can designate an authorized agent to make a request under the CCPA on your behalf. We may deny a request from an authorized agent that does not submit proof that they have been validly authorized to act on your behalf in accordance with the CCPA. - You may request to opt out from future selling or sharing of your personal information to third parties. Upon receiving an opt-out request, we will act upon the request as soon as feasibly possible, but no later than fifteen (15) days from the date of the request submission. To exercise these rights, you can contact us by email at dataleadsfuture@gmail.com, or by referring to the contact details at the bottom of this document. If you have a complaint about how we handle your data, we would like to hear from you. **14\. DO VIRGINIA RESIDENTS HAVE SPECIFIC PRIVACY RIGHTS?** **In Short:* Yes, if you are a resident of Virginia, you may be granted specific rights regarding access to and use of your personal information.* **Virginia CDPA Privacy Notice** Under the Virginia Consumer Data Protection Act (CDPA): "Consumer" means a natural person who is a resident of the Commonwealth acting only in an individual or household context. It does not include a natural person acting in a commercial or employment context. "Personal data" means any information that is linked or reasonably linkable to an identified or identifiable natural person. "Personal data" does not include de-identified data or publicly available information. "Sale of personal data" means the exchange of personal data for monetary consideration. If this definition "consumer" applies to you, we must adhere to certain rights and obligations regarding your personal data. The information we collect, use, and disclose about you will vary depending on how you interact with us and our Services. To find out more, please visit the following links: - [Personal data we collect](#infocollect) - [How we use your personal data](#infouse) - [When and with whom we share your personal data](#whoshare) Your rights with respect to your personal data - Right to be informed whether or not we are processing your personal data - Right to access your personal data - Right to correct inaccuracies in your personal data - Right to request deletion of your personal data - Right to obtain a copy of the personal data you previously shared with us - Right to opt out of the processing of your personal data if it is used for targeted advertising, the sale of personal data, or profiling in furtherance of decisions that produce legal or similarly significant effects ("profiling") We have not sold any personal data to third parties for business or commercial purposes. We will not sell personal data in the future belonging to website visitors, users, and other consumers. Exercise your rights provided under the Virginia CDPA More information about our data collection and sharing practices can be found in this privacy notice and our Cookie Notice: [http://www.dataleadsfuture.com/cookies](https://www.dataleadsfuture.com/cookies). You may contact us by email at dataleadsfuture@gmail.com, by submitting a [data subject access request](https://app.termly.io/notify/489d8eac-3a1d-4368-bc83-aa3cb2203283?ref=dataleadsfuture.com), or by referring to the contact details at the bottom of this document. If you are using an authorized agent to exercise your rights, we may deny a request if the authorized agent does not submit proof that they have been validly authorized to act on your behalf. Verification process We may request that you provide additional information reasonably necessary to verify you and your consumer's request. If you submit the request through an authorized agent, we may need to collect additional information to verify your identity before processing your request. Upon receiving your request, we will respond without undue delay, but in all cases, within forty-five (45) days of receipt. The response period may be extended once by forty-five (45) additional days when reasonably necessary. We will inform you of any such extension within the initial 45-day response period, together with the reason for the extension. Right to appeal If we decline to take action regarding your request, we will inform you of our decision and reasoning behind it. If you wish to appeal our decision, please email us at dataleadsfuture@gmail.com. Within sixty (60) days of receipt of an appeal, we will inform you in writing of any action taken or not taken in response to the appeal, including a written explanation of the reasons for the decisions. If your appeal if denied, you may contact the [Attorney General to submit a complaint](https://www.oag.state.va.us/consumer-protection/index.php/file-a-complaint?ref=dataleadsfuture.com). **15\. DO WE MAKE UPDATES TO THIS NOTICE?** **In Short:* Yes, we will update this notice as necessary to stay compliant with relevant laws.* We may update this privacy notice from time to time. The updated version will be indicated by an updated "Revised" date and the updated version will be effective as soon as it is accessible. If we make material changes to this privacy notice, we may notify you either by prominently posting a notice of such changes or by directly sending you a notification. We encourage you to review this privacy notice frequently to be informed of how we are protecting your information. **16\. HOW CAN YOU CONTACT US ABOUT THIS NOTICE?** If you have questions or comments about this notice, you may email us at dataleadsfuture@gmail.com or contact us by post at: Data Leads Future \_\_\_\_\_\_\_\_\_\_ dataleadsfuture@gmail.com \_\_\_\_\_\_\_\_\_\_ Hong Kong **17\. HOW CAN YOU REVIEW, UPDATE, OR DELETE THE DATA WE COLLECT FROM YOU?** Based on the applicable laws of your country, you may have the right to request access to the personal information we collect from you, change that information, or delete it. To request to review, update, or delete your personal information, please fill out and submit a [data subject access request](https://app.termly.io/notify/489d8eac-3a1d-4368-bc83-aa3cb2203283?ref=dataleadsfuture.com). This privacy policy was created using Termly's [Privacy Policy Generator](https://termly.io/products/privacy-policy-generator/?ref=dataleadsfuture.com). ### Cookies URL: https://www.dataleadsfuture.com/cookies/ Last updated: 2023-08-27T06:01:46.000Z **COOKIE POLICY** **Last updated August 27, 2023** This Cookie Policy explains how Data Leads Future ("**Company**," "**we**," "**us**," and "**our**") uses cookies and similar technologies to recognize you when you visit our website at [https://www.dataleadsfuture.com](https://www.dataleadsfuture.com/) ("**Website**"). It explains what these technologies are and why we use them, as well as your rights to control our use of them. In some cases we may use cookies to collect personal information, or that becomes personal information if we combine it with other information. **What are cookies?** Cookies are small data files that are placed on your computer or mobile device when you visit a website. Cookies are widely used by website owners in order to make their websites work, or to work more efficiently, as well as to provide reporting information. Cookies set by the website owner (in this case, Data Leads Future) are called "first-party cookies." Cookies set by parties other than the website owner are called "third-party cookies." Third-party cookies enable third-party features or functionality to be provided on or through the website (e.g., advertising, interactive content, and analytics). The parties that set these third-party cookies can recognize your computer both when it visits the website in question and also when it visits certain other websites. **Why do we use cookies?** We use first- and third-party cookies for several reasons. Some cookies are required for technical reasons in order for our Website to operate, and we refer to these as "essential" or "strictly necessary" cookies. Other cookies also enable us to track and target the interests of our users to enhance the experience on our Online Properties. Third parties serve cookies through our Website for advertising, analytics, and other purposes. This is described in more detail below. **How can I control cookies?** You have the right to decide whether to accept or reject cookies. You can exercise your cookie rights by setting your preferences in the Cookie Consent Manager. The Cookie Consent Manager allows you to select which categories of cookies you accept or reject. Essential cookies cannot be rejected as they are strictly necessary to provide you with services. The Cookie Consent Manager can be found in the notification banner and on our website. If you choose to reject cookies, you may still use our website though your access to some functionality and areas of our website may be restricted. You may also set or amend your web browser controls to accept or refuse cookies. The specific types of first- and third-party cookies served through our Website and the purposes they perform are described in the table below (please note that the specific cookies served may vary depending on the specific Online Properties you visit): **Essential website cookies:** These cookies are strictly necessary to provide you with services available through our Website and to use some of its features, such as access to secure areas. | Name: | m | | ----------- | ---------------------------------------------------------------------------------------- | | Purpose: | Tracks the user's session for Stripe | | Provider: | m.stripe.com | | Service: | Stripe [View Service Privacy Policy](https://stripe.com/privacy?ref=dataleadsfuture.com) | | Country: | United States | | Type: | server\_cookie | | Expires in: | 1 year 11 months 29 days | | Name: | \_\_stripe\_mid | | ----------- | ---------------------------------------------------------------------------------------------- | | Purpose: | Fraud prevention and detection | | Provider: | .www.dataleadsfuture.com | | Service: | Stripe [View Service Privacy Policy](https://stripe.com/en-nl/privacy?ref=dataleadsfuture.com) | | Country: | Sweden | | Type: | http\_cookie | | Expires in: | 11 months 30 days | | Name: | \_\_stripe\_sid | | ----------- | ---------------------------------------------------------------------------------------------- | | Purpose: | Fraud prevention and detection | | Provider: | .www.dataleadsfuture.com | | Service: | Stripe [View Service Privacy Policy](https://stripe.com/en-nl/privacy?ref=dataleadsfuture.com) | | Country: | Sweden | | Type: | http\_cookie | | Expires in: | 30 minutes | **Analytics and customization cookies:** These cookies collect information that is used either in aggregate form to help us understand how our Website is being used or how effective our marketing campaigns are, or to help us customize our Website for you. | Name: | \_ga | | ----------- | ----------------------------------------------------------------------------------------------------------- | | Purpose: | Records a particular ID used to come up with data about website usage by the user | | Provider: | .dataleadsfuture.com | | Service: | Google Analytics [View Service Privacy Policy](https://policies.google.com/privacy?ref=dataleadsfuture.com) | | Country: | Netherlands | | Type: | http\_cookie | | Expires in: | 1 year 1 month 4 days | | Name: | \_ga\_# | | ----------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- | | Purpose: | Used to distinguish individual users by means of designation of a randomly generated number as client identifier, which allows calculation of visits and sessions | | Provider: | .dataleadsfuture.com | | Service: | Google Analytics [View Service Privacy Policy](https://policies.google.com/privacy?ref=dataleadsfuture.com) | | Country: | Netherlands | | Type: | http\_cookie | | Expires in: | 1 year 1 month 4 days | **Unclassified cookies:** These are cookies that have not yet been categorized. We are in the process of classifying these cookies with the help of their providers. | Name: | ghost-history | | ----------- | ----------------------- | | Purpose: | \_\_\_\_\_\_\_\_\_\_ | | Provider: | www.dataleadsfuture.com | | Service: | \_\_\_\_\_\_\_\_\_\_ | | Country: | Sweden | | Type: | html\_local\_storage | | Expires in: | persistent | **How can I control cookies on my browser?** As the means by which you can refuse cookies through your web browser controls vary from browser to browser, you should visit your browser's help menu for more information. The following is information about how to manage cookies on the most popular browsers: - [Chrome](https://support.google.com/chrome/answer/95647?ref=dataleadsfuture.com#zippy=%2Callow-or-block-cookies) - [Internet Explorer](https://support.microsoft.com/en-us/windows/delete-and-manage-cookies-168dab11-0753-043d-7c16-ede5947fc64d?ref=dataleadsfuture.com) - [Firefox](https://support.mozilla.org/en-US/kb/enhanced-tracking-protection-firefox-desktop?redirectslug=enable-and-disable-cookies-website-preferences&redirectlocale=en-US&ref=dataleadsfuture.com) - [Safari](https://support.apple.com/en-ie/guide/safari/sfri11471/mac?ref=dataleadsfuture.com) - [Edge](https://support.microsoft.com/en-us/windows/microsoft-edge-browsing-data-and-privacy-bb8174ba-9d73-dcf2-9b4a-c582b4e640dd?ref=dataleadsfuture.com) - [Opera](https://help.opera.com/en/latest/web-preferences/?ref=dataleadsfuture.com) In addition, most advertising networks offer you a way to opt out of targeted advertising. If you would like to find out more information, please visit: - [Digital Advertising Alliance](http://www.aboutads.info/choices/?ref=dataleadsfuture.com) - [Digital Advertising Alliance of Canada](https://youradchoices.ca/?ref=dataleadsfuture.com) - [European Interactive Digital Advertising Alliance](http://www.youronlinechoices.com/?ref=dataleadsfuture.com) **What about other tracking technologies, like web beacons?** Cookies are not the only way to recognize or track visitors to a website. We may use other, similar technologies from time to time, like web beacons (sometimes called "tracking pixels" or "clear gifs"). These are tiny graphics files that contain a unique identifier that enables us to recognize when someone has visited our Website or opened an email including them. This allows us, for example, to monitor the traffic patterns of users from one page within a website to another, to deliver or communicate with cookies, to understand whether you have come to the website from an online advertisement displayed on a third-party website, to improve site performance, and to measure the success of email marketing campaigns. In many instances, these technologies are reliant on cookies to function properly, and so declining cookies will impair their functioning. **Do you use Flash cookies or Local Shared Objects?** Websites may also use so-called "Flash Cookies" (also known as Local Shared Objects or "LSOs") to, among other things, collect and store information about your use of our services, fraud prevention, and for other site operations. If you do not want Flash Cookies stored on your computer, you can adjust the settings of your Flash player to block Flash Cookies storage using the tools contained in the [Website Storage Settings Panel](http://www.macromedia.com/support/documentation/en/flashplayer/help/settings%5Fmanager07.html?ref=dataleadsfuture.com). You can also control Flash Cookies by going to the [Global Storage Settings Panel](http://www.macromedia.com/support/documentation/en/flashplayer/help/settings%5Fmanager03.html?ref=dataleadsfuture.com) and following the instructions (which may include instructions that explain, for example, how to delete existing Flash Cookies (referred to "information" on the Macromedia site), how to prevent Flash LSOs from being placed on your computer without your being asked, and (for Flash Player 8 and later) how to block Flash Cookies that are not being delivered by the operator of the page you are on at the time). Please note that setting the Flash Player to restrict or limit acceptance of Flash Cookies may reduce or impede the functionality of some Flash applications, including, potentially, Flash applications used in connection with our services or online content. **Do you serve targeted advertising?** Third parties may serve cookies on your computer or mobile device to serve advertising through our Website. These companies may use information about your visits to this and other websites in order to provide relevant advertisements about goods and services that you may be interested in. They may also employ technology that is used to measure the effectiveness of advertisements. They can accomplish this by using cookies or web beacons to collect information about your visits to this and other sites in order to provide relevant advertisements about goods and services of potential interest to you. The information collected through this process does not enable us or them to identify your name, contact details, or other details that directly identify you unless you choose to provide these. **How often will you update this Cookie Policy?** We may update this Cookie Policy from time to time in order to reflect, for example, changes to the cookies we use or for other operational, legal, or regulatory reasons. Please therefore revisit this Cookie Policy regularly to stay informed about our use of cookies and related technologies. The date at the top of this Cookie Policy indicates when it was last updated. **Where can I get further information?** If you have any questions about our use of cookies or other technologies, please email us at dataleadsfuture@gmail.com or by post to: Data Leads Future \_\_\_\_\_\_\_\_\_\_ \_\_\_\_\_\_\_\_\_\_ United States This cookie policy was created using Termly's [Cookie Consent Manager](https://termly.io/products/cookie-consent-manager/?ref=dataleadsfuture.com). ### Disclaimer for Data Leads Future URL: https://www.dataleadsfuture.com/disclaimer/ Last updated: 2026-07-01T15:17:10.000Z Please read this disclaimer carefully before using the Data Leads Future website (the "Website") operated by Peng Qian ("us", "we", "our"). ## General Information Purposes Only The information provided by Data Leads Future is for general informational purposes only. All information on the website is provided in good faith, however, we make no representation or warranty of any kind, express or implied, regarding the accuracy, adequacy, validity, reliability, availability, or completeness of any information on the website. ## Not Professional Advice The information presented on Data Leads Future is not professional advice and should not be treated as such. The content is provided for general informational purposes only and is not a substitute for professional advice. Accordingly, before taking any actions based upon such information, we encourage you to consult with the appropriate professionals. ## No Endorsements The website may contain links to other websites or content belonging to or originating from third parties or links to websites and features. Such external links are not investigated, monitored, or checked for accuracy, adequacy, validity, reliability, availability, or completeness by us. WE DO NOT WARRANT, ENDORSE, GUARANTEE, OR ASSUME RESPONSIBILITY FOR THE ACCURACY OR RELIABILITY OF ANY INFORMATION OFFERED BY THIRD-PARTY WEBSITES LINKED THROUGH THE SITE. ## Personal Responsibility You must take personal responsibility for the use of our website and its content. By using the website, you agree to take full responsibility for any harm or damage you suffer as a result of the use or reliance on any content provided on the website. ## No Liability In no event shall we be liable for any indirect, punitive, incidental, special, consequential damages, or any damages whatsoever, including, without limitation, damages for loss of use, data, or profits, arising out of or in any way connected with the use or performance of the website, the delay or inability to use the website, the provision of or failure to provide services, or for any content, software, products, and services obtained through the website, or otherwise arising out of the use of the website. ## Affiliate Disclosure Data Leads Future may have financial relationships with some of the merchants mentioned on the website and may be compensated if consumers choose to click on the links located throughout the content on the site and generate sales for the said merchant. And it won't cost you anything extra. Data Leads Future is in alliance partnership with the following brands: - Coursera ## Investment Risks Any investment information provided on the website may carry an inherent risk, and the value of investments may go down as well as up. You may want to consider seeking advice from a financial and investment advisor before making any investment. ## Consulting Services Any consulting services provided by Data Leads Future to paid members are intended for general guidance on matters of interest only. The application and impact of laws can vary widely based on the specific facts involved. Given the changing nature of laws, rules, and regulations, there may be delays, omissions, or inaccuracies in information contained in the consulting services. ## Governing Law and Jurisdiction The website is operated under the laws of the United States and the European Union. If you access the website from outside the United States or the European Union, you do so at your own risk and are responsible for compliance with the laws of your jurisdiction. ## Consent By using our website, you hereby consent to our disclaimer and agree to its terms. ## Changes to This Disclaimer Should we update, amend, or make any changes to this document, those changes will be prominently posted here. If you require any more information or have any questions about our site's disclaimer, please feel free to contact us by email at dataleadsfuture@gmail.com. This Disclaimer was created for Data Leads Future on January 10, 2024. Peng Qian The Author, Founder of Data Leads Future ### Terms and Services for Data Leads Future URL: https://www.dataleadsfuture.com/terms-and-services/ Last updated: 2024-01-10T07:35:25.000Z ## Introduction Welcome to Data Leads Future ([https://www.dataleadsfuture.com](https://www.dataleadsfuture.com/)), a blog dedicated to providing educational content on data science, including tutorials, practical applications, coding techniques, machine learning, and algorithmic trading in stock and currency markets. This Terms and Services agreement outlines the rules and regulations for the use of Data Leads Future's website and services. By accessing and using this website, you agree to comply with the following terms and conditions. If you disagree with any part of these terms, please do not use our website. ## Collection of Information Data Leads Future collects personal information such as email addresses to send subscribers the latest articles and updates. The collection and use of personal information are governed by our Privacy Policy, which complies with the General Data Protection Regulation (GDPR) for users from the European Union and applicable U.S. privacy laws. ## Monetization Methods Data Leads Future operates on a monetization model that may include but is not limited to advertisements, affiliate marketing, sponsored content, email marketing, and subscription-based memberships. We ensure all monetization efforts are transparent and comply with relevant advertising standards and regulations. ## Consulting Services Data Leads Future may offer consulting services to paid members. These services provide general guidance and are not a substitute for professional advice. The user assumes responsibility for the use of any advice or guidance given. ## Copyright Compliance We respect intellectual property rights and insist that the content hosted on the website does not infringe upon the copyright or intellectual property of others. Any use of copyrighted material is done under the provision of "fair use." ## User Responsibilities Users of Data Leads Future are expected to respect the integrity of the website and its content. Unauthorized use, distribution, or reproduction of content provided by Data Leads Future is strictly prohibited and may result in legal action. ## Limitation of Liability Data Leads Future and its operators will not be liable for any indirect, incidental, consequential, special, exemplary, or punitive damages arising out of or related to the use of the website. This limitation of liability applies to the fullest extent permitted by law and survives any termination or expiration of this agreement or your use of Data Leads Future or the services found on this website. ## Governing Law These Terms shall be governed and construed in accordance with the laws of the United States and the European Union, without regard to its conflict of law provisions. Our failure to enforce any right or provision of these Terms will not be considered a waiver of those rights. ## Amendments to Terms and Services Data Leads Future reserves the right to modify these terms at any time. By continuing to access or use our service after any revisions become effective, you agree to be bound by the revised terms. ## Contact Us For any questions regarding these Terms and Services, please contact us at dataleadsfuture@gmail.com. ## Acknowledgment By using our website, you acknowledge that you have read and understand these Terms and Services and agree to be bound by them. ## Effective Date These terms were last updated on January 10, 2024. This Terms and Services page has been created to ensure the responsible and compliant operation of the Data Leads Future blog. Our aim is to foster a trusted environment where our content can be enjoyed without legal concerns. Peng Qian The Author, Founder of Data Leads Future ### Contact URL: https://www.dataleadsfuture.com/contact/ Last updated: 2024-04-19T06:18:01.000Z You can find me on: [Twitter](https://twitter.com/qtalen?ref=dataleadsfuture.com), [LinkedIn](https://www.linkedin.com/in/qtalen/?ref=dataleadsfuture.com), [Facebook](https://www.facebook.com/qtalen). If you have any questions about Data Leads Future, you can contact me at [dataleadsfuture@gmail.com](mailto:dataleadsfuture@gmail.com). ### Tags URL: https://www.dataleadsfuture.com/tags/ Last updated: 2026-05-29T08:54:54.000Z [Harness Engineering - Data Leads FutureOne practical story every month, sharing my hard-learned experiences in the enterprise AI space. Every article saves you 40+ hours of work.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-2330856173653bac7b53b40d6d9e0f745402472ecf68c7b8efd0bba1fb3d0321.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_1-8aa0c3b06c75de81597ae443e56e9788fa609b04b607939c5f7953edc29b23a7.webp)](https://www.dataleadsfuture.com/tag/harness-engineering/) [Agentic AI - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-18.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic_ai_2.webp)](https://www.dataleadsfuture.com/tag/agentic-ai/) [Generative AI - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-14.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/generative_ai_1.webp)](https://www.dataleadsfuture.com/tag/generative-ai/) [Microsoft Agent Framework - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-59.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/maf.webp)](https://www.dataleadsfuture.com/tag/microsoft-agent-framework/) [AutoGen - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-31.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/autogen.webp)](https://www.dataleadsfuture.com/tag/autogen/) [Llama Index - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-15.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/llama_index.webp)](https://www.dataleadsfuture.com/tag/llama-index/) [Chainlit - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-20.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Chainlit.webp)](https://www.dataleadsfuture.com/tag/chainlit/) [Machine Learning - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-16.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/machine_learning.webp)](https://www.dataleadsfuture.com/tag/machine-learning/) [Python Craft - Data Leads FutureRefine your Python projects, focus on the latest technological advancements, and infuse your code with spirit.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-17.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/python_craft.webp)](https://www.dataleadsfuture.com/tag/python-craft/) [Career Growth - Data Leads FutureStay ahead in the world of AI and data science! Whether you’re a beginner or an expert, we provide cutting-edge, practical insights to help you solve real-world problems and achieve your goals.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-19.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/career_growth_1-1.webp)](https://www.dataleadsfuture.com/tag/career-growth/) ## Posts ### How I Cut Kimi K3 Costs in OpenCode URL: https://www.dataleadsfuture.com/how-i-cut-kimi-k3-costs-in-opencode/ Last updated: 2026-08-05T07:35:38.000Z *Disclaimer: This post contains affiliate links. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* ## Introduction Last week, I burned through 7 days' worth of quota in a single day while using the `Kimi K3` model on Kimi's Allegreto plan. `Kimi K3` is genuinely great. Its performance is on par with `Claude Fable 5` and `GPT 5.6 SOL`, so I ended up going full throttle with it all day long. But the cost is ridiculous. It runs way higher than `GLM-5.2`, `DeepSeek-V4`, or even `Kimi K2.7`. For someone like me who got spoiled by the cheap API of `DeepSeek-V4`, that was just not okay. So I started optimizing how I use OpenCode. The goal was to do more with the same `Kimi K3` quota while keeping quality about the same. After a few days of work, the results are pretty solid. The Allegreto plan now covers a full week of development for me. No more sitting around two days out of five waiting for the weekly `Kimi K3` quota to reset. If these methods work for me, they should work for you too. So this article is a quick write-up of what I've done, and I hope it helps you lower your `Kimi K3` costs in OpenCode. All the source code mentioned in this article is at the bottom. Feel free to grab it. --- ## Pick the Right Provider The most fundamental way to cut costs is picking the right provider. The official Coding Plan is the best option. Based on various reports, the third and fourth tiers of the Coding Plan offer dozens of times more value per dollar than the API at the same price. That's a great deal. Besides the official Coding Plan, if you'd rather pay per use or call a third-party API, you can just use open platforms that support Kimi K3, like [**Novita.ai**](https://fas.st/t/cBJ7iuNS?ref=dataleadsfuture.com) (which offers ultra-low-latency endpoints and generous free credits for new users) and OpenRouter. The setup and how it works are the same. Since I already purchased the official Allegreto plan upfront, I'll use the Coding Plan models as the primary example throughout this article. --- ## Pick the Right Model and Thinking Level (Variants) ### Pick the right model The Coding Plan gives you access to two models: K2.7 Code and K3\. K3 requires Allegreto or above, and only the K3 model supports 1M context length. This week, Kimi also released a K3 model with 256K context, with the model ID `k3-256k`. At the same time, the official docs confirmed that the standard K3 model consumes twice the quota of `k3-256k`. That's probably why I burned through a whole week's quota right out of the gate. Without a second thought, I switched my default model to `k3-256k`. ![A comparison of model capabilities provided on the official website.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-13.png) A comparison of model capabilities provided on the official website. Screenshot from [Kimi](https://www.kimi.com/code/docs/en/kimi-code/models.html?ref=dataleadsfuture.com) You might wonder: since I'm used to `DeepSeek-V4`'s 1M context, will 256K be enough? I don't think you need to worry too much about that. The 1M context mainly helps your cache hit rate stay high as conversations get longer. But context rot is still a real problem as the context grows. So if you've ever felt like `DeepSeek-V4` gets dumber after a long session, your instinct is right. That's context rot at work. On top of that, we normally use frameworks like OpenSpec for SDD (Spec-Driven Development). All the plans and specs worked out earlier with the model get saved as files on disk. Whether you use the `/compact` command or start a new session, the model reads context from those files. The message history doesn't need to be that long. Then there are situations where you need to scan a large codebase or pull in a lot of information from the web. For those cases, we use sub-agents running in separate sub-sessions to handle the research, then return only the key findings to the main session. That approach cuts down context usage a lot. All things considered, 256K context is plenty for now. For the `Plan` and `Build` agents, just use `k3-256k` directly. If you're using the API from [**Novita.ai**](https://fas.st/t/cBJ7iuNS?ref=dataleadsfuture.com), it's even simpler. Just use `Kimi K3` straight up. ### Pick the thinking level For a long time, Kimi models felt slow. That's because before K3, Kimi didn't support the `reasoning_effort` parameter. Every call defaulted to maximum thinking, so each request took forever to finish. The K3 release added support for `reasoning_effort`, with three levels: `low`, `high`, and `max`. But when the model first launched last week, only `max` was available. That meant every call generated massive thinking tokens through a long chain-of-thought process, which burned through a huge amount of token budget. Good news: starting this week, both K3 models support `low` and `high`. If you're setting up `Kimi K3` in OpenCode for the first time this week, the default thinking level is `High`. If you configured `Kimi K3` last week, make sure you change the thinking level from `Max` to `High`. If you care more about code quality than cost, or you're doing complex research and don't want lower thinking intensity to hurt your results, there's a middle ground. Use `Max` thinking in the `Plan` agent for architecture planning, then use `High` thinking in the `Build` agent for code execution. One thing to watch out for here is that, according to the [official docs](https://www.kimi.com/code/docs/en/kimi-code/models.html?ref=dataleadsfuture.com), switching the `reasoning_effort` value invalidates the context cache. Right after the switch, the model immediately refills the cache using your existing message history, which costs extra tokens. The official recommendation is to open a new session before switching `reasoning_effort`, so you avoid paying to refill old messages into the cache. There's a more elegant solution though. Set up a dedicated sub-agent with `Max` thinking, specifically for architecture decisions and hard problems. I'll cover that in the next section. --- ## Use Sub-Agents to Build an Efficient Team With the model and thinking level choices from earlier, you should already be saving a good chunk of tokens. If you want to cut costs even further, the best move is to use sub-agents, each with its own model and prompt, so each agent fits the task it handles. Why does this work? If you've ever led a dev team, you know that not everyone needs to be a superstar engineer, because superstars come with superstar price tags. From a cost and output perspective, the smartest call is matching the right task to the right person. A team runs best that way. So which sub-agents should you set up to get more out of `Kimi K3` per dollar? ### Explore Have you ever noticed that the bigger your codebase, the faster your token usage spikes? That comes down to how LLM caching and codebase retrieval work together. OpenCode puts tools, skills, system prompts, and other mostly static content at the top of the message list when it sends requests to the model. The model caches all of that, so subsequent prompts don't cost much extra. The real token hog is actually your most recent prompt. Before making any code changes, OpenCode uses `grep` to search through relevant code. When the codebase is large, even a few `grep` calls can pull in a massive amount of text. That's where most of your tokens go. ![The biggest token consumers are user messages and tool call results.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-14.png) The biggest token consumers are user messages and tool call results. Image by Author When you need to pull information from multiple files or web sources, doing that in the main session makes things worse. A huge amount of irrelevant text gets stuffed into the main message list, blowing up both the model's context window and your quota. That's why OpenCode uses an `explore` sub-agent to handle complex research in a separate sub-session. Once the research is done, only the conclusions get returned to the main session. That keeps the main session's token usage lean. By default, `explore` has no model assigned to it, so it falls back to whatever model the primary agent uses, which is `Kimi K3`. That keeps costs high. The fix is to give the `explore` agent its own dedicated model. I strongly recommend `deepseek-v4-flash`. The `flash` model is cheap, has a long context window, and is plenty capable for text retrieval. Setup is simple. Just add a few lines to `~/.config/opencode/opencode.json`: ```json { "$schema": "https://opencode.ai/config.json", "agent": { "explore": { "model": "novita-ai/deepseek/deepseek-v4-flash" } } } ``` **Quick setup tip:** If you don't have a Novita account yet, grab your API Key from [**Novita AI**](https://fas.st/t/cBJ7iuNS?ref=dataleadsfuture.com), then paste it into the API KEY field under the Novita provider. ![Move the retrieval cost from the main session to a sub-session running on the DeepSeek Flash model.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-15.png) Move the retrieval cost from the main session to a sub-session running on the DeepSeek Flash model. Image by Author Give it a try. When your bill drops significantly, you'll thank me. ### Executor With the `explore` sub-agent and its dedicated model, we've already cut a big chunk of input token costs. Now let's look at how to reduce output token costs. Output token costs mainly come from actually writing code. There's a common misconception here. You might think that since `Kimi K3` is so capable, it must be the one writing the final code to make sure things work. That's not actually true. Writing code is a lot like a skilled programmer looking at a design doc and typing out what's on the screen. Most of it runs on muscle memory. If the coding task is well-planned in advance, and the model doing the actual output isn't too far behind in capability, then whether `Kimi K3` or `deepseek-v4-flash` writes the code, the quality difference is not that big. What really matters is who thinks. Good system design and knowing exactly what code will solve the problem efficiently, that's the important part. So we can keep the `Plan` primary agent for task planning and detailed spec generation (I recommend OpenSpec for SDD planning here), then hand off the actual coding to an `executor` sub-agent. This agent doesn't need to think. It just needs to write the code. OpenCode doesn't come with an `executor` sub-agent out of the box, so we build one ourselves. You don't have to agonize over the prompt. Just open OpenCode and tell `Kimi K3` what you want, like this: ```Markdown I want to create a sub-agent called `executor`. Its main job is to complete coding and file editing tasks quickly at low cost. It doesn't do any research or information retrieval. That's the primary agent's job. Its only responsibility is to implement code fast based on the primary agent's plan. ``` OpenCode will write an `executor` agent file under `~/.config/opencode/agents/` based on what you described. Once it's done, restart OpenCode, and you're good to go. You can also give the `executor` more responsibilities. For example, you can allow it to run concurrently to speed things up: ```Markdown This agent should support concurrent work. When the tasks to be implemented have no dependencies on each other, multiple `executor` agents (no more than 5) can be called in parallel to finish the work faster. ``` Don't forget to set the `executor` agent's model to `deepseek-v4-flash`. ```Markdown --- description: ... mode: subagent model: novita-ai/deepseek/deepseek-v4-flash permission: webfetch: deny websearch: deny task: deny skill: deny --- ``` Restart OpenCode again. If you're not sure whether `executor` is configured correctly, ask OpenCode to run a self-check and evaluate whether the `executor` is working the way you designed it. ![Ask OpenCode to run a self-check to see if the executor is working properly.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-16.png) Ask OpenCode to run a self-check to see if the `executor` is working properly. Image by Author OpenCode will generate some temporary coding tasks and hand them to `executor` to complete. It will also try to ask `executor` to do things it's not permitted to do, just to see if it follows the rules. When `executor` finishes, it returns a summary to the primary agent about whether the task was completed successfully. `Kimi K3` never has to output the actual code, so those expensive output token costs stay low. At the same time, `deepseek-v4-flash` is capable enough that your code quality won't take a noticeable hit. One important thing to keep in mind with the `executor` agent: don't let the `deepseek-v4` model write frontend code. Due to gaps in post-training, `deepseek-v4` really struggles with frontend visual design. So you should restrict the `executor` agent from handling frontend work: ```Markdown This agent does not support frontend development. Any frontend-related work should never be assigned to this agent. ``` Or you can set the `executor` agent to still use the `Kimi K3` model, but with the variant set to `low`. If you mostly write backend code, `deepseek-v4-flash` is enough. If you write frontend code regularly, I'd go with the K3 model plus `variant=low`. That brings up a natural question: how do we make the most out of the thinking level parameter? ### Architect To keep `Kimi K3` fast and reduce token consumption during the thinking process, we set the default thinking level to `High`. I mentioned earlier that you can set the `Plan` agent to `Max` thinking during the planning phase. But what if you need K3 to hit Fable 5-level performance to crack a hard problem? You'd have to set `reasoning_effort` to `Max`. The official docs say that switching `reasoning_effort` invalidates the existing context cache, so the recommendation is to open a new session first. But that's not realistic in practice. The times when you actually need `reasoning_effort=max` are usually when you're stuck on an architecture tradeoff or a bug you can't fix after several tries. At that point, your message list is already packed with critical context. Starting a new session means K3 has to re-scan all the relevant code from scratch, and it's doing that scan at `max` thinking intensity, which burns even more tokens. And then if you want to switch back from `max` to `high`, the cache has to be refilled all over again, wasting another round of tokens. No matter how you slice it, it's not worth it. So why does a sub-agent solve this? Because we can set up a K3 sub-agent with `reasoning_effort` locked at `max` (I call this one `architect`). When a hard problem needs solving, the primary agent first summarizes its current session into a condensed brief, then sends that brief along with the problem to the `architect` agent through a fresh sub-session. At that point, `architect` already understands the full context, so it doesn't need to re-scan all the documents the way a brand-new session would. It only needs to look up a few extra details it cares about, then focus entirely on thinking through the solution. I also turn off edit permissions for the `architect` agent. That way K3 isn't burning `max`\-level thinking on actually writing code fixes. It just figures out the solution. The primary agent handles everything else. Building this sub-agent is straightforward. Just open OpenCode and ask it to write the prompt for you: ```Markdown Design an architect sub-agent specifically for making architecture decisions, debugging hard bugs, and doing code reviews. ``` Don't forget to assign `architect` the Kimi K3 model with `variant` set to `max`: ```Markdown --- description: ... mode: subagent model: novita-ai/moonshotai/kimi-k3 variant: max permission: edit: deny task: deny --- ``` To stop the primary agent from getting lazy and routing every problem to `architect`, add a guardrail in the sub-agent description. Require the primary agent to try at least two times on its own before delegating: ```Markdown Hard bugs must go through at least two independent hypothesis-and-verify rounds by the primary agent with no resolution before they can be delegated here. ``` After restarting OpenCode, ask it to run a self-check to verify that `architect` applies the right thinking level when tackling hard problems: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-17.png) Run a self-check on the availability of the `architect` agent. Image by Author ### General Last but not least, don't overlook the `general` sub-agent. As one of OpenCode's built-in sub-agents, this one has always had a confusing identity. Looking at the source code, it's meant for complex research and retrieval work, which puts it in similar territory as `architect`, only to be called when heavy reasoning is needed. But it also supports concurrent calls to speed up research, which makes its token burn rate closer to `executor`. The role feels split in two. This agent rarely gets called during normal use. But if you ignore it, one day a complex deep research task might come up, and it'll get called concurrently to run complex web searches and research work, all using the default `Kimi K3 High`. So if you want tight control over your LLM token usage, make sure you assign this sub-agent a cheaper model. In my view, `deepseek-v4-pro` is a great fit here. It's capable enough for complex research, and the token cost isn't much higher than `deepseek-v4-flash`: ```json { "$schema": "https://opencode.ai/config.json", "agent": { "explore": { "model": "novita-ai/deepseek/deepseek-v4-flash" }, "general": { "model": "novita-ai/deepseek/deepseek-v4-pro" } } } ``` --- ## Conclusion That covers everything I've done to use Kimi K3 at lower cost inside OpenCode. With Kimi K3's model weights now publicly released, these methods come at the perfect time to help you experience this excellent model without breaking the bank. In my experience, after applying all these settings, even during high-intensity development work, my Allegretto plan went from "work two days, rest five" to "work five days, rest two." Even if your main models are Claude or GPT, these optimization approaches still apply. On top of that, the sub-agent setup I recommended here will noticeably speed up your development workflow and boost your overall efficiency. That's all for today. I'm [Mr. Qian](https://www.linkedin.com/in/qtalen?ref=dataleadsfuture.com), and I focus on enterprise-level AI Agent applications in practice. If you have any questions about this article, feel free to leave a comment, and I'll get back to you as soon as I can. Thank you for reading and subscribing. Feel free to share this article with your friends so more people can benefit from it. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further Reading How I built a coding workflow in OpenCode, Oh-My-OpenCode-Slim, and OpenSpec that rivals Claude Code: [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-e5c5bb35-cbaa-434b-baea-547f937998e0.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-59d35e77-6277-4b4b-8e8f-23f411b077fb.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) By adding a reflective agent to the OpenSpec workflow, I managed to get DeepSeek-V4-Pro to perform at the level of Opus: [Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec WorkflowStop letting bad spec files tank your code quality![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-5a85b02b-2380-40c1-8b28-50ee54deefbb.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/9f04a900-aa9a-4fd2-bd70-6f3fe3783129-f9ad9e37-a235-4e82-b193-90540c1d7319.webp)](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) The concept of Loop Engineering has been getting a lot of buzz lately, so I decided to give it a shot in OpenCode. The results were surprisingly good: [No Plugins Needed, I Built a Fully Automated Coding Loop in OpenCodeUsing DeepSeek-V4 for low-cost Loop Engineering![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-abc0b774-ec22-4b1b-8119-0037d9dcc229.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-3-comprass-b8cec65e-5414-414c-80bf-e2b49730c256.webp)](https://www.dataleadsfuture.com/no-plugins-needed-i-built-a-fully-automated-coding-loop-in-opencode/) --- Here's the source code for this article. Sign up now to get it for free. [Grab the Source Code ](#/portal/signup) ### No Plugins Needed, I Built a Fully Automated Coding Loop in OpenCode URL: https://www.dataleadsfuture.com/no-plugins-needed-i-built-a-fully-automated-coding-loop-in-opencode/ Last updated: 2026-08-05T07:34:46.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* This article was last updated on July 27, 2026. Today I want to talk about how I built an agentic coding loop inside OpenCode. People call this Loop Engineering. To show you how well this coding loop holds up, I used `DeepSeek-V4-Pro` and `DeepSeek-V4-Flash` for every example in this article. It turns out that with the right design, DeepSeek models can pull off high-quality coding loops at a very low cost. All the source code in this article is available at the end of the post. Let's get into it. --- ## Introduction If you've been following the AI agent space recently, you've probably heard the term Loop Engineering. Claude Code and OpenClaw both mention it. But nobody really explains what it means. Until Andrew Ng posted [a clear breakdown of what Loop Engineering actually is.](https://www.deeplearning.ai/the-batch/issue-359?ref=dataleadsfuture.com) That post gave me the direction I needed to build it inside OpenCode. --- ## What Andrew Ng Says About the Coding Loop Andrew Ng's explanation centers around this diagram: ![3 key product development loops.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-1.png) 3 key product development loops. Image by [Andrew Ng](https://www.deeplearning.ai/the-batch/issue-359?ref=dataleadsfuture.com) Three loops, simply put: - The first loop is the code agent's implementation loop. You give the agent a product spec and a measurable target. The agent starts building features on its own, runs tests, and keeps iterating until every requirement in the spec is met. - The second loop is the engineer feedback loop. Here the engineer acts as QA for their own product. They test what the coding agent built, and check if it matches their vision. If something is off, they write a new spec and kick off another round in the first loop. - The third loop is the external feedback loop. Once the engineer is happy with the product, it goes to the open-source community or gets handed to a product team. Real user feedback comes in. The engineer collects that feedback regularly, feeds it back into the engineer feedback loop, and from there back into the code implementation loop. All three loops keep running. With AI in the mix, they push the product forward until it gets to where it needs to be. --- ## My Take on the Coding Loop The way I see it, these three loops map to two commands in Claude or Codex: `/goal` and `/loop`. `/goal` covers the first and second loops. Once a user writes a product requirements doc or a spec, they pass the task to the agent through the `/goal` command. The agent iterates on its own, runs its own tests, and opens a Pull Request. Then the engineer reviews the code and tests it manually. If everything looks good, they merge the PR and ship. Next comes the external feedback loop. You can use the `/loop` command to set up scheduled tasks that pull issues from GitHub or Jira on a timer. The agent turns those issues into requirement docs, and `/goal` takes it from there. ![Use /goal to create a new task, and use /loop to schedule it to run on a timer.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-2.png) Use `/goal` to create a new task, and use `/loop` to schedule it to run on a timer. Image by Author There is a gap in the traditional understanding here. When a user calls `/goal`, they just pass in what they want done. How it gets done, and what the exit condition is, often gets left up to a strong reasoning model like Claude Opus or GPT 5.5 to figure out. But in a proper coding loop, you need to give the agent a product spec and a measurable completion target so it knows when to exit. Something like this: ```Markdown /goal Use the spec I wrote to build this chess game. Use Playwright MCP for end-to-end testing. Verify the game supports castling, pawn promotion, checkmate, stalemate, and rating calculation. Fix any bugs found during testing and re-run. Loop no more than 20 times. ``` That is a solid starting instruction for a coding loop. It needs a goal, a verification method, and a max iteration count. And you have to write something like that every single time, either right after the `/goal` command or inside the spec itself. In my OpenCode version, I am not going that route. I just want to tell `/goal` what I want. The agent handles everything else. The rest of this article shows you how I built that enhanced `/goal` loop. (The `/loop` command is a separate topic I will cover in a future post.) --- If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. --- ## See the Final Result First Before getting into the implementation, let me show you what the `/goal` command actually does in a few scenarios, from simple to complex. ### 1\. Write a Fibonacci script This test checks whether the agent picks the best algorithm. The prompt is: ```Markdown /goal Write a Fibonacci calculation script with the best possible performance. ``` ![What's pretty cool is that DeepSeek went with a fast doubling algorithm, and it's incredibly performant.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-3.png) What's pretty cool is that DeepSeek went with a fast doubling algorithm, and it's incredibly performant. Image by Author ### 2\. Build a Tower of Hanoi web game Everyone knows this game. You could honestly just prompt an LLM directly and get it built. But that predictability is exactly what makes it useful for testing whether each step of the coding loop is working correctly. The prompt is: ```Markdown /goal Build a playable Tower of Hanoi web game. ``` ![The Towers of Hanoi game built using the /goal command comes with a built-in AI solver.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/hanoi_tower.gif) The Towers of Hanoi game built using the `/goal` command comes with a built-in AI solver. Image by Author ### 3\. Build a chess web game Maybe you are not impressed by the Tower of Hanoi example, since a regular prompt to any frontier model can do the same thing without a coding loop. Fair. The chess game experiment is something you should actually try for yourself. The prompt is: ```Markdown /goal Build a playable chess web game. Include easy, medium, and hard difficulty levels. No online multiplayer needed. Use a Python backend as the AI engine. ``` ![A full-stack chess game with both frontend and backend, built using the /goal command. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/web_chess.gif) A full-stack chess game with both frontend and backend, built using the `/goal` command. Image by Author In this experiment, OpenCode first does a quick requirements check with the user after receiving the task. It offers suggestions along the way so you can just click next. Once the full requirements list is confirmed, it enters autonomous mode. ![The /goal command breaks down your goal into a list of requirements and then implements them one by one.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-6.png) The `/goal` command breaks down your goal into a list of requirements and then implements them one by one. Image by Author It handles architecture design, builds each feature module, runs unit tests, edge case tests, and end-to-end tests. It only hands the result back to the user when everything passes. Pretty cool, right? Want to know how it works? Let's go. --- ## Building the Coding Loop ### Overall design OpenCode updates very fast. Almost every release breaks something or introduces small bugs. So I am not building yet another OpenCode plugin, since a plugin can easily break after an update. The good news is that OpenCode lets you define commands and agents using plain Markdown files. Write a few lines of prompts, save the file as Markdown, and drop it in the right directory. That mechanism is the foundation for building a powerful coding loop on the cheap. Let me walk you through the interaction flow of the whole loop. When a user runs `/goal ` to start a new coding task, OpenCode sends that task to an agent I call the orchestrator. In this setup, that agent is `goal-orch`. The orchestrator's job is to talk to the user and review completed work. When it receives a task, it does not start coding right away. It first has a short conversation with the user to clarify the details, then breaks the task down into a list of atomic requirements. This is different from OpenCode's built-in `Plan` agent. The `/goal` orchestrator does not ask open-ended questions. It takes its best guess at what the user wants and asks the user to confirm whether those sub-requirements are correct. This works much better when the user is not a software engineer. It is like a capable assistant breaking down tasks for their manager without making the manager feel lost. Once the requirements list is confirmed, the orchestrator saves it as a `reqs-manifest.md` file. This file tracks requirement status and serves as the acceptance record throughout the loop. After that, the orchestrator enters autonomous mode. It starts an internal loop that calls a worker agent called `goal-worker` and sends each requirement to it. `goal-worker` handles code implementation. Before writing any code, it does architecture planning first. If there are technology choices involved, it raises questions. It does not ask the user though. It asks `goal-orch`. Once `goal-orch` answers, `goal-worker` continues. After the architecture plan is done, `goal-worker` enters the loop and starts implementing each requirement. Each requirement includes feature code, edge case handling, and unit test code. When all of that is done, `goal-worker` sends an implementation report back to `goal-orch`. The orchestrator checks the report against the requirement description, confirms completion, and updates the status in the requirements list. Inside the loop, `goal-orch` also builds a DAG from the full requirements list to map out dependencies. Requirements with no dependencies between them get executed in parallel with multiple `goal-worker` instances, which speeds things up significantly. Once all requirements are done, `goal-orch` runs acceptance tests against the list, checks unit test coverage, and for tasks involving frontend work, it launches end-to-end tests using `browser use` to confirm everything works. When everything passes, the orchestrator writes an implementation report and notifies the user that the task is complete. Then it waits for the next task. The execution sequence looks like this: ![The sequence diagram for how the /goal command runs.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-7.png) The sequence diagram for how the `/goal` command runs. Image by Author Now, let me walk through how each piece is built. ### Implementing the`/goal`command Since we want `/goal` to be the entry point for the coding loop, just like in Claude Code, which is the first thing to build in OpenCode. Thankfully, creating a command in OpenCode is pretty simple. Just drop a Markdown command definition file into the `.opencode/commands/` directory. Two things to keep in mind when writing the Markdown command file: 1. In the `frontmatter`, configure a separate agent through the `agent` parameter. Without a dedicated agent, OpenCode sends the command straight to the default agent. The default agent has no Loop Engineering prompts, so the command fails. For `/goal`, I set up an orchestrator agent named `goal-orch`, which I will cover next. 2. In the body of the command, use the `$ARGUMENTS` placeholder to represent the task the user passes when calling the command. When the command sends the message to the orchestrator, this placeholder gets replaced with whatever the user typed, and the full assembled message goes to the orchestrator. ### Implementing the`goal-orch`agent Now let's build the `goal-orch` agent, which acts as the orchestrator. It handles user communication, requirements clarification, kicks off the development loop, and calls `goal-worker`, and verifies that work is complete. Since we have built OpenCode agents multiple times in previous articles, I will not go into every detail here. Just a few things worth noting: 1. Model choice. `goal-orch` handles workflow orchestration overall, so it needs strong reasoning. I am a fan of the DeepSeek-V4 series, so I went with `deepseek-v4-pro`. 2. System prompt. The prompt needs to cover both workflow orchestration and the agent's own tasks, so it is on the longer side. But that is fine. When you use `/goal` to send a task, the `goal-orch` system prompt automatically replaces OpenCode's default prompt, so there is no need to worry about excessive token usage. 3. `goal-orch` must be a primary agent. OpenCode has two agent types: primary agents and subagents. The key difference is that primary agents can call subagents to delegate work, but subagents cannot. Since `goal-orch` needs to call `goal-worker` inside the loop, so it has to be a primary agent. The only downside is that it shows up in the agent list at the bottom of the OpenCode interface, but you get used to it. 4. The interaction mechanism between primary and subagents. During the loop, when `goal-worker` hits a tricky problem or needs to make a technology choice, it does not ask the user. It returns the question to `goal-orch` through a session return. This requires `goal-orch` to answer the question and then continue the previous sub-session. OpenCode uses a `task` tool to send messages to subagents. This tool accepts a `task_id` parameter. The first call to the `task` tool opens a new subagent session, and the return value includes a `task_id`. After `goal-orch` answers `goal-worker`'s question, it passes the answer and the previous `task_id` back into the `task` tool. This resumes the task in the original sub-session and preserves context. 5. Put simply, it works like a foreman and a worker. The foreman sends the worker off to do a job. The worker gives the foreman a reference number. If the worker hits a problem, they come back and ask the foreman. The foreman solves the problem, takes the reference number and the answer back to the worker, and the worker picks up exactly where they left off. No starting over. ![OpenCode uses task_id to keep track of the context in sub-sessions.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-8.png) OpenCode uses `task_id` to keep track of the context in sub-sessions. Image by Author 1. Parallel task execution. Even though we are building a fully automatic coding loop, the orchestrator can still run multiple `goal-worker` instances in parallel for requirements that have no dependencies on each other. This speeds up the overall task significantly. After `goal-orch` confirms the requirements list with the user, it does not jump straight into the coding loop. It first maps out the dependency relationships between requirements into a DAG. When it finds requirements with no dependencies between them, it kicks off parallel execution to finish the task faster. ![The requirements are recalculated into a DAG based on their dependencies, and then executed concurrently.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/image-9.png) The requirements are recalculated into a DAG based on their dependencies, and then executed concurrently. Image by Author 1. Every completed task gets documented. A coding loop handles a large volume of requirements at once, and this process can run for hours. As the task keeps going, the context grows longer and the agent's ability to follow the prompt gets weaker. At that point, we cannot rely on session messages alone to track the status of each requirement. The orchestrator will start missing key steps in the workflow. So when designing `goal-orch`, I included a requirement in the prompt: subagents must not only update the requirement status in `reqs-manifest.md` after each step, they also generate a documentation note for each completed requirement. That way, once the coding loop ends, `goal-orch` can review those documents to see which tasks finished cleanly and which ones hit blockers that need human attention. 2. End-to-end testing with `browser use`. Since we are going fully automatic, why not go all the way? When `goal-orch` detects that the task involves a frontend page, it calls the `agent-browser` skill or `@playwright-mcp` during the final review phase to run end-to-end tests, and captures screenshots as proof that the feature works. Since I am using `deepseek-v4` models, I also call the `@observer` agent for image reading when needed. I covered the `@observer` agent implementation in a previous article. That covers the notable design decisions behind `goal-orch`. I am not going to paste the full prompt here since prompts are easy to generate with an LLM once you understand the principles. Grab the full source code at the end of the article. #### It's Loop Engineering, and it's also Graph Engineering When `goal-orch` calculates the dependencies across all requirements before starting implementation and compiles them into a DAG, our `/goal` goes beyond Loop Engineering in the narrow sense. It levels up into a new concept: Graph Engineering. In Graph Engineering, the agent no longer completes tasks one by one in a straight line. It pre-arranges all tasks into a dependency map, then lets `goal-worker` move through that map. `goal-orch` only checks `goal-worker`'s results at the end of the map. ![From Loop Engineering to Graph Engineering.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/opencode_loop-graph.drawio-1.png) From Loop Engineering to Graph Engineering. Image by Author Now let's look at `goal-worker`. ### Implementing the`goal-worker`agent Compared to `goal-orch`, `goal-worker` is much simpler by design. It just does the work. The simpler its instructions, the better it performs. That said, I did set a few expectations for it in this coding loop: 1. Model choice. For tasks that do not need deep reasoning, `deepseek-v4-flash` honestly outperforms `pro`. From my experience, in situations involving function calling and skill loading, `flash` is more accurate than `pro` and costs much less. So for `goal-worker`, which is all about execution, `deepseek-v4-flash` is the right call. 2. Ask before acting. Even though it is a subagent, I gave `goal-worker` the ability to ask questions. Whether it is before starting or in the middle of a task, whenever it hits a technical question, `goal-worker` stops and asks `goal-orch` for input. This makes full use of the stronger reasoning model higher up in the chain. 3. Plan first, then act. Just like the outer loop driven by `goal-orch`, `goal-worker` follows the same rule: plan before doing. Before implementing any new requirement, `goal-worker` writes out a plan that covers the tech stack, what it intends to change, and how it will handle unit tests and edge cases. That plan goes to `goal-orch` for review. Only after `goal-orch` signs off does `goal-worker` start coding. This mechanism maximizes the chance of a successful implementation. 4. Use modern package managers. This is not strictly required, but in practice I noticed that DeepSeek tends to default to traditional tools like `pip` and `npm` for global dependency installs. As a developer, that is something I can not live with. So I explicitly require `goal-worker` to use modern package managers like `uv` and `pnpm`. This also keeps your environment clean when the coding loop is building something for you. That is the full implementation plan for my OpenCode coding loop. I did not paste source code into the article body since all of it is at the end. Go grab it there. I also tried to explain the thinking behind this coding loop and how to build it clearly enough that you could just drop this article into OpenCode and have it reconstruct a working `/goal` command for you. From there, tweak it however you want and build your own version of Loop Engineering. In the age of AI, what is really impossible anymore? Before wrapping up, let me quickly cover how to actually use this coding loop. --- ## How to Use This Coding Loop First, go to the end of this article and open the GitHub link I shared. Then come back here and keep reading. The source code for the coding loop lives in its own git repo. The only directory you need to care about is `.opencode/`. To use this coding loop across all your OpenCode projects, copy the `agents` and `commands` directories from that folder into `~/.config/opencode/`, then restart OpenCode. If you only want to try it in one specific project, just copy the `.opencode/` directory into your project root and restart OpenCode. If you just want to take it for a quick spin, clone the repo locally, create a new branch, and run `opencode` from that branch to start the environment. After that, type the `/goal` command in OpenCode with whatever task you want the coding loop to handle. For example, if you want to build the typing practice game Andrew Ng mentioned in his post, just enter this: ```Markdown /goal I want to build a typing practice web app for kids. At the bottom of the screen are 9 keys representing the positions of 10 fingers, with the thumbs sharing a wider spacebar key that takes up two slots. Above those 9 keys is a gradient-transparent rectangle. Letters fall down from the top toward that rectangle, aligned to their matching key positions. Press the right key as the letter enters the zone and it counts as a hit, triggering a hit effect. Consecutive correct hits build up increasingly dramatic effects. The falling speed gradually increases. At the top of the screen is a scoreboard that adds 1 point for each hit, with a pop animation. The whole thing should feel exciting and dopamine-triggering, with plenty of visual encouragement. ``` ![A typing mini-game created with /goal command.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/typing-game.gif) A typing mini-game created with `/goal` command. Image by Author You can be as detailed or as vague as you want. Either way, once you hit enter, the coding loop starts running. `goal-orch` will open a conversation with you to clarify the requirements. You can fill in more details as the conversation develops, or just choose `recommend` for everything and go make yourself a coffee while it does the work. --- ## Conclusion That is the story of how I built this coding loop inside OpenCode. Before I wrap up, let me answer a few questions you might have. **1\. Can open-source models support this coding loop?** Yes, I think so. In this article, I used `deepseek-v4-pro` for `goal-orch` and `deepseek-v4-flash` for `goal-worker`, and both worked well. With the strict constraints of the loop workflow in place, they can handle most of your tasks automatically. Also, DeepSeek is releasing the official version of DeepSeek-V4 in mid-July. It will almost certainly come with post-training improvements, and I expect even better results from it. **2\. What results should I expect?** The coding loop will spend hours working through your task, but do not set your expectations too high for what a single `/goal` run can produce. That is true even with the most powerful Claude or GPT models. As Andrew Ng said, the code implementation loop is just the first of three loops. You need to step into the second loop, the engineer feedback loop. Try out what `/goal` produced, find all the bugs and things that do not match your vision, organize those into a new spec, and kick off another coding loop. Keep iterating. That is how the product gets better over time. **3\. Can the coding loop replace SDD frameworks like OpenSpec?** No. They serve different scenarios. The coding loop is great for product managers or team leads who want to quickly spin up a minimum viable product from market insights or their own ideas. It prioritizes speed and automation. It is for people who want to hand things off and let the agent run on its own. But if your product is going to production, if you have high standards for code quality and implementation details, if you have a detailed product spec, and you are an experienced software engineer who knows exactly what the product should look like, then OpenSpec or another SDD plugin is the right tool. Start from a solid foundation and build your full product step by step. I will keep updating this article as I use the loop more, and the Q&A section at the end will keep growing. If you have any questions, leave a comment, and I will get back to you as soon as I can. Thanks for reading and subscribing. Next time you are sitting there sipping coffee while OpenCode codes away on its own, if this article comes to mind, please share it with a friend. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further Reading How I built a coding workflow in OpenCode, Oh-My-OpenCode-Slim, and OpenSpec that rivals Claude Code: [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-8ff27605-9172-4dc0-bdb3-360fcc53b0bc.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-b8bc482d-63d1-4b72-bce8-32c2d73114fe.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) How I added a reflection agent to the OpenSpec workflow and got DeepSeek-V4 to match or beat Claude Opus: [Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec WorkflowStop letting bad spec files tank your code quality![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-8e41de04-bfe2-4d86-97d2-91a39a8470f2.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/9f04a900-aa9a-4fd2-bd70-6f3fe3783129-4c821d9d-2e0b-4652-8d7f-fdef160cab5f.webp)](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) Still can't get your DeepSeek-V4 or GLM-5.2 to read images? Try my method: [DeepSeek-V4 Can’t Read Images? I Made It ReadDon’t wait for a multimodal model, you can use it now![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-c5d8144e-d3c4-4d74-b8f0-4b6d4deec490.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-bc0420ff-3660-45b9-9c1c-14bdcfeb9967.webp)](https://www.dataleadsfuture.com/deepseek-v4-cant-read-images-i-made-it-read/) --- ## Source Code Here's the source code for this article. Sign up now to get it for free. [Grab the Source Code ](#/portal/signup) ### Why Well-Written Prompts Are Almost Always Structured URL: https://www.dataleadsfuture.com/why-well-written-prompts-are-almost-always-structured/ Last updated: 2026-07-27T01:40:05.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* People are arguing again about whether prompts should use HTML or Markdown. Last year, the debate was Markdown vs. JSON. But to me, the answer has nothing to do with the format itself. Markdown, JSON, HTML, they all work well for the same reason: structure. As for which one to pick, it depends on which format appeared more frequently in your model's pretraining data. Prompt engineering is not about feeding the model more information. It's about helping the model and the human put limited attention in the right place. Structured language does exactly that. With just a few markup symbols, it pre-allocates attention across the text. Through Markdown, JSON, and HTML tags, language stops being a linear stream of words that only flows left to right. It becomes something with hierarchy, emphasis, attention weights, and a multi-dimensional shape. Structured text does not just change how an LLM pays attention. It changes how you pay attention, too. Let's look at a passage in plain language first: > *We're going on a picnic on July 1st, and the destination is Central Park. We're driving. You take the ring road from the south side of the city, and I'll drive straight down from the north. We'll meet at Entrance 2 of the park at 8 AM. Remember to bring tuna sandwiches, bottled water, and a Bluetooth speaker. I'll bring the picnic blanket, the tent, chicken sandwiches, and bug spray. Also, don't forget to book the park tickets in advance on the app. You can only book one day ahead, so search for Central Park in the app.* Now let's try structuring that same text, and strip out the natural language filler (conjunctions, prepositions, all the connective tissue, since the structure already shows the relationships and none of that is needed anymore). Here's what it looks like after Markdown rendering: ```Markdown ### Picnic Location Central Park ### Meeting Time and Place - July 1st, meet at 8 AM - **Entrance 2** of the park ### Getting There - You: drive via the ring road - Me: drive straight from the north side ### What to Bring - You: **tuna** sandwiches, bottled water, Bluetooth speaker - Me: **chicken** sandwiches, picnic blanket, tent, bug spray ### Important Reminder - Must book in advance on the app, search *Central Park* - **Can only book one day ahead** ``` Now, if someone asks which entrance to meet at, you can find the answer so much faster, right? --- If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. --- Through categorization and emphasis, the important information already pulls your attention toward it. It doesn't get buried inside a wall of text that means nothing at first glance. All of that change comes from tweaking just a handful of symbols to introduce a complete structured syntax: ```Markdown ### Heading **Bold** *Italic* - List 1. Ordered list ``` So back to the original point: why do structured prompts work better? Because structured text is no longer natural language. Natural language lets attention drift, for both LLMs and humans. Structured language anchors attention on purpose. A lot of decisions and responses depend on where attention lands, not on how smart you are. Here's a simple example. You need to buy a bag of pasta at the grocery store. You're incredibly smart, but you're hungry right now. Your attention goes straight to the serving suggestion photo on the packaging, and it looks delicious. So you grab a bag without thinking twice. Standing next to you is a nutritionist. Whose attention goes straight to the ingredient list and nutrition facts. His reaction: too many calories, I'm trying to lose weight, putting it back. You didn't miss the calorie count because you can't read a nutrition label. You missed it because the pretty food photo on the package caught your eye first. But if someone told you beforehand that high-carb foods make weight gain easy, you'd probably think twice next time you reach for pasta. That story makes two points: 1. Attention drives what happens next. 2. Attention can be shaped, at least to some degree. Guiding attention inside a prompt matters a lot, then. Rather than hoping an LLM will notice the ingredient list on its own, just call it out directly in the prompt: ```Markdown ### Pay Close Attention!! 1. **Do not** focus only on the cover image. 2. **Always check the ingredient list**. ``` You mark it because you're not leaving it to chance. That's the simplest logic behind structured prompts, and it's logic both you and an LLM can follow. So back to the question at the start: HTML or Markdown? It doesn't matter. What matters is whether you're using structure to anchor attention for the model and for yourself. The symbols are just the tool. Structure is the whole point. Thanks for reading and subscribing. If you found this helpful, feel free to share it with a friend. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further Reading Here are prompts and LLM tips I often use at work, sharing them with you: [Share My LLM Prompts and Tips That Make Work and Learning Super EfficientCan they help you too?![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-35c2653b-02fc-447e-a397-832556e882a8.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-1-9f708a23-7644-4d08-ae3f-bf88853cbe2f.webp)](https://www.dataleadsfuture.com/share-my-llm-prompts-and-tips-that-make-work-and-learning-super-efficient/) You pour your heart and soul into building an AI agent system, only to find out after launch that it's just another barely functional demo product. Sound familiar? [My Agent System Looks Powerful but Is Just Industrial TrashSingle agent, low robustness, and position bias![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-bcfe2178-230f-4652-a228-a2918b8ccae2.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-1-a158ded8-1380-4526-a390-d468b19b654c.webp)](https://www.dataleadsfuture.com/my-agent-system-looks-powerful-but-is-just-industrial-trash/) ### DeepSeek-V4 Can't Read Images? I Made It Read URL: https://www.dataleadsfuture.com/deepseek-v4-cant-read-images-i-made-it-read/ Last updated: 2026-08-05T07:30:40.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* ## Introduction I gotta say, DeepSeek-V4 is cheap, works great, and has a long context window. It's already my go-to coding model. The new GLM-5.2 model is also pretty impressive. But as of mid-June, the multimodal versions of these models still haven't been released. That means anything involving images like reading error screenshots, interpreting charts, or recreating pages from visual mockups is totally out of their reach. However, I found a workaround: I built a little plugin called `observer` in OpenCode. By letting DeepSeek-V4 call a multimodal agent, it can indirectly get the ability to "see" images. After polishing it for over a month, this plugin now handles all my image-related coding tasks at work. Today, I'm sharing how I made it, hoping it might help you too. The plugin code and agent definitions mentioned in this article are at the end. Feel free to grab them. --- ## Demo of Real-World Results Before diving into the long tutorial, you probably care most about how well this plugin works and whether it is worth your time to try. So let me show you some screenshots of the plugin in action. ### 1\. Interpreting error stack traces We start with the simplest task: have `deepseek-v4` interpret a screenshot of an error stack trace and find key information. I randomly picked a screenshot of an error I encountered at work: ![A screenshot of a common error stack.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-1.png) A screenshot of a common error stack. Image by Author Then in OpenCode Desktop, I sent this image to the `plan` agent using `deepseek-v4-pro` and asked it to provide a solution: ![The deepseek-v4-pro agent quickly picked up the error message.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-2.png) The deepseek-v4-pro agent quickly picked up the error message. Image by Author As you can see, the `plan` agent gave an answer based on the screenshot information. ### 2\. Interpreting charts Another multimodal use case is interpreting charts from documents. For this example, I took a screenshot of a company's annual revenue chart and tested it. I still used the `plan` agent with `deepseek-v4-pro`. For an extra challenge, I asked the agent to give some key insights on the numbers in the chart: ![A screenshot of a listed company's financial report.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-3.png) A screenshot of a listed company's financial report. Image by AlphaStreet The agent read the numbers from the chart and provided some key insights: ![The agent accurately spotted the data in the chart and offered key insights.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-4.png) The agent accurately spotted the data in the chart and offered key insights. Image by Author ### 3\. Developing HTML pages from designs In frontend development, the biggest demand for multimodal capability is recreating visual designs. Here I found a design with complex page elements to see if the `build` agent using `deepseek-v4-flash` could recreate the page: ![A screenshot of a web design draft.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-5.png) A screenshot of a web design draft. Image by dribbble.com Here is the recreated page: ![The page that deepseek v4 flash recreated.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-6.png) The page that deepseek v4 flash recreated. Image by Author One thing is sure: the `deepseek-v4-flash` model generated the frontend code, and it only took one prompt to get this result. It did not get a 100% match, but with a few more rounds of conversation, you can tweak it until it is perfect. Keep in mind `deepseek-v4-flash` is dirt cheap. It costs several times or even ten times less than multimodal models like `kimi k2.6` or `qwen3.7 plus`. They are not in the same league. Of course, you can also crop a section of the page, mark the areas that need attention, and ask DeepSeek to adjust them, like this: ![You can take a screenshot of the webpage and have deepseek v4 flash make adjustments. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-8.png) You can take a screenshot of the webpage and have deepseek-v4-flash make adjustments. Image by Author The agent perceives the marked area and gives the primary agent an adjustment plan per your request. ### 4\. Generating HTML pages from hand-drawn sketches Maybe you are like me and have zero design skills. No problem. We can hand-draw rough sketches. The agent can understand them. For example, in a recent project, I hand-drew a few web page design sketches: ![This is a hand-drawn sketch of the webpage.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-9.png) This is a hand-drawn sketch of the webpage. Image by Author Then `deepseek-v4-flash` helped me recreate the page: ![The agent restored the page based on my handwritten reference.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/image-11.png) The agent restored the page based on my handwritten reference. Image by Author Impressive, right? If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. --- ## Detailed Implementation Walkthrough I know you cannot wait any longer. Let me jump straight into the implementation details. The whole image-reading plugin consists of two parts: - A sub-agent configured with a multimodal LLM. It runs in a separate sub-session, reads the images uploaded by the user, parses them into detailed text descriptions based on the scenario, and returns the results to the DeepSeek model in the main session. - An OpenCode plugin that intercepts images uploaded by the user, saves them as files, and triggers the sub-agent to read the images at the right time. In other words, the plugin is the "dispatcher," and the sub-agent is the "image reader." They work together through independent sub-sessions without messing up the main session's context. Let me start with the design of the agent. ### Designing the image-reading agent Since the source code is at the end of this article, I won't paste it here. I will only cover the design thinking behind the agent. This agent does the actual image reading, so make sure it uses a multimodal LLM. Here I used the `kimi-for-coding/k2p6` model. The setup is simple. Just put the configuration in the frontmatter of the agent's Markdown file (that's the YAML block wrapped in three dashes at the very top of the file). Of course, it won't match a native multimodal agent. This approach, where another multimodal agent converts an image to a text description and then passes it back to DeepSeek, inevitably loses a lot of information. To capture as many image details as possible, I broke the reading process into different scenarios. Each scenario corresponds to a different working mode, with its own trigger keywords and output format: **Mode A:** Page Restoration. Keywords like: restore, HTML, page, design mockup, etc. Main task: describe the image at the pixel level with precision to help the main agent write an identical HTML page: ```Markdown ### Mode A: Page Restoration **Signal words**: Reproduce, HTML, page, design mockup, screenshot reproduction, refactor, frontend, CSS, layout, slice images, implement, pixel-perfect, 1:1, precise reproduction, replicate, mobile, app screenshot, component, visual design, Figma, XD **Task**: Describe the webpage/app interface screenshot with pixel-level precision, helping the main agent write an identical HTML/CSS page. **Simplified mode**: If the signal words contain one of `rough`, `approximate`, `briefly describe`, `quick and simple`, only output A1 (page overview) + A5 (page text list), skip the rest of the sections. ``` To describe the page layout well, I also told the agent to output the page structure using ASCII art. My experiments show this ASCII approach is effective. **Mode B:** Issue Location and Fix. When my screenshot has areas marked with red boxes, arrows, etc., and I ask the agent to pay special attention, this mode kicks in. ```Markdown ### Mode B: Issue Location and Fix **Signal words**: issue, fix, adjust, wrong, error, bug, tweak, something off, not normal, mark, red box, arrow, circle, look here, this area, this part, skewed, misaligned, spacing, not aligned, wrong color, wrong font, overflow, overlap **Task**: Identify the problem areas marked or pointed out in the screenshot, analyze the symptoms and possible causes, and give specific fix suggestions. ``` **Mode C:** Error Log Extraction. I use this mode a lot in daily work. For example, when a remote computer blocks the clipboard, we can take a screenshot and let the agent analyze the error stack trace in the image. ```Markdown ### Mode C: Error Log Extraction **Signal words**: error, log, error, stack, stack trace, exception, exception, crash, traceback, warning, warning, fail, crash, 500, 404, timeout, panic, fail **Task**: Extract the error/log text from the screenshot precisely, word for word, keeping all technical details so the main agent can locate and fix the code. ``` **Mode D:** Text/Conversation Extraction and Analysis. This is the basic OCR function. Just recognize the conversation roles, text hierarchy, and content relationships. ```Markdown ### Mode D: Text/Conversation Extraction and Analysis (Default) **Signal words**: extract text, OCR, recognize text, read text, conversation, copywriting, clarify, content relationships, what was said, transcribe, organize **Task**: Extract all text from the image, clarify conversation roles, text hierarchy, content relationships, and logical structure. ``` **Mode E:** Chart Interpretation. This mode extracts data metrics from charts, as I demonstrated earlier. ```Markdown ### Mode E: Chart/Data Visualization Extraction **Signal words**: charts, line charts, bar charts, pie charts, scatter plots, radar charts, heatmaps, area charts, trends, data visualization **Task**: Extract data points, axis labels, and trend info from chart screenshots for the main agent to analyze data relationships. ``` I did not decide on these five modes all at once. I added them one by one as I needed them in real scenarios. This brings up a problem: several modes share overlapping signal words, but each mode has a different output format. If there is a conflict, which mode should the agent use to read the image? Here I used a priority method. I defined five priority levels as follows: `C (Error Log Extraction) > E (Chart Data Extraction) > B (Issue Location and Fix) > A (Page Restoration) > D (Text/Conversation Extraction and Analysis)` If I add new modes later, I will adjust the priorities. Finally, I set Mode D as the default. When none of the previous modes match the user's request, it will use OCR mode to handle the image reading task. Once the image-reading agent is designed, just drop the Markdown file into the `~/.config/opencode/agents/` directory and it will take effect. ### Designing the OpenCode plugin Compared to the agent, the plugin design is much simpler. It is just a single JS file with a little over 100 lines, and it handles two main jobs: Use the `experimental.chat.system.transform` hook. Every time a request is sent to the model, it checks if the model has multimodal capability. If not, it adds this prompt to the system prompt: ```Markdown ## Image Reading - You should use the @observer sub-agent to read images. - When a message like [Image saved to: ] appears in the conversation, call the @observer sub-agent and tell it to read the image file at that path. ``` Use the `chat.message` hook. When the user pastes an image into the input box and sends a message, the image exists as base64-encoded text in the message body. The plugin intercepts this data, decodes it, and saves it as an image file in a temp directory. Then it replaces the original image content in the message with text like `[Image saved to: ]`. The plugin's entire workflow looks like this: ![The workflow of plugin.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/06/deepseek-image.ver2.drawio.png) The workflow of plugin. Image by Author You need to place the plugin in the `~/.config/opencode/plugin/` directory for it to work. That covers the implementation principles of the OpenCode plugin and the agent that add image-reading capability to `deepseek-v4`. --- ## Why Did I Design This Plugin? After reading this whole practice, you might have a question: why don't I just start a new conversation and use a multimodal model to read images directly? Why go through all this trouble? What is the benefit? I plan to answer your question from the following angles: - Cost. We use `deepseek-v4`'s two models for their huge cost-effectiveness. Switching to a multimodal model directly would defeat the purpose of using them. - Keep the conversation context. For example, interpreting an error stack trace or extracting text from a conversation screenshot. These actions happen during a coding task. The existing conversation context is crucial for finding a fix based on the error. If you start a new conversation to read the image, you lose that context. - Different models support different context lengths. The `deepseek-v4` model supports a 1M context length, while models like `kimi k2.6` only support 256k. If you temporarily switch to k2.6 in the middle of a conversation to read an image, it might immediately trigger OpenCode's context compression, causing key info loss. So the best solution is to use a sub-agent that reads the image in a separate sub-session and returns the result to the main agent. --- ## Conclusion We still wait, excited for DeepSeek to release a multimodal model. Once native multimodal capability arrives, we can pack up all this fussing with plugins and sub-agents. But until then, the `observer` plugin is a workable solution. And the process itself has value: once you understand OpenCode's plugin and agent orchestration, you can apply the same thinking to patch capability gaps in other models. This mechanism is not limited to OpenCode. It works on other coding agent platforms too, just adjust the implementation a bit. The plugin code and agent definitions mentioned in this article are all at the end. Grab them there. If you think this article helps you, feel free to share it with your friends. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further Reading Give my programming workflow built with OpenCode, OMO-Slim, and OpenSpec a try? [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-2f3b5557-139f-47a1-9227-b0bed7d7a406.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-7d04fb67-3c58-41fb-828d-f3375518399d.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) In the previous article, I added a reflection mechanism to the OpenSpec workflow, making `deepseek-v4-pro` match or even exceed the coding performance of the `opus` model. Click here to learn more: [Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec WorkflowStop letting bad spec files tank your code quality![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-a22fe7fa-1fac-45fd-8c27-0588b8b44168.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/9f04a900-aa9a-4fd2-bd70-6f3fe3783129-b317c55f-4779-4d95-9a54-d57caa4eed7b.webp)](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) Better code quality, but spend less money. [How I Cut Kimi K3 Costs in OpenCodeBetter code quality, but spend less money![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-fb16fe3f-229e-43a3-a52f-09c128aaf373.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover_tiny-4ca67667-4a82-4180-939e-8f956d65f1a7.webp)](https://www.dataleadsfuture.com/how-i-cut-kimi-k3-costs-in-opencode/) --- ## Source Code Here's the source code for this article. Sign up now to get it for free. [Grab the Source Code ](#/portal/signup) ### Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec Workflow URL: https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/ Last updated: 2026-08-05T07:30:15.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* ## Introduction In this article, I want to walk you through how I introduced a reflection mechanism into the OpenSpec workflow inside OpenCode, and how that dramatically improved the quality of AI-generated code. After nearly a month of testing, this reflection workflow has gotten DeepSeek-V4-pro in OpenCode to perform at roughly the same level as Claude Opus 4.6\. The only cost is some extra review time and a few more tokens. Trust me, it's worth it. Can't wait to find out how? Let's get into it. --- ## Why This Works I've been using AI coding for a while now. Compared to Claude Code, I prefer building my own SDD-based coding workflow with OpenCode and OpenSpec. I wrote a well-received article specifically about this OpenCode workflow: [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://static.ghost.org/v5.0.0/images/link-icon.svg)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-5a74fbb6366cd99192a7d0a8c5a77c96965c01527c3bec8db395e0700f93e9f9.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) With OpenSpec, and the explore → propose → apply → verify → archive workflow loop, we can finally get LLMs to handle complex project development. ![With the OpenSpec workflow loop, LLMs can finally write usable code.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image.png) With the OpenSpec workflow loop, LLMs can finally write usable code. Image by Author But just like you've probably run into, even with the SDD workflow, no matter which model I use, GPT 5.5 or the latest DeepSeek-V4-pro, the AI still inevitably produces hidden bugs or piles up messy code. Code I'd never feel comfortable putting in production. My first fix was to call a `@reviewer` sub-agent after the `/opsx-apply` phase to do a code review on the changes. Sometimes that worked and caught architectural or implementation issues. But the impact was limited. Often I'd only discover something was wrong after using the project for a while: a scenario wasn't covered, edge cases were missed, or one part of the code got updated but a related module didn't. Later I stepped back and looked at the whole AI coding workflow again, and that's when I spotted the real problem. As programmers, we always focus on whether our code is good, so we naturally look at things from the code level. But we completely overlooked the quality of the proposal files that OpenSpec generates. We'd finish discussing requirements, generate the proposal file, and then just let the AI start implementing. That's how input-level bugs get introduced. When things go wrong, we blame the model. Think about it, back in the traditional coding era, when a product manager handed over a requirements doc, there was a critical step before we started coding: requirements review. We wouldn't touch the keyboard until every issue in the requirements doc or design doc was sorted out. So why did we forget this step in the AI coding era? That's exactly what we're going to fix today: add the requirements review step back into the OpenSpec workflow and see if it makes AI-generated code better. --- ## How to Do It The SDD workflow isn't anything exotic. If you're familiar with multi-agent system design patterns, you'll recognize that SDD is basically the `plan-execute` pattern. One agent breaks the user's task into a step-by-step plan file, then another agent follows that plan to execute the task. This lets the agent system handle complex work. ![The Plan agent makes the plan, the Executor agent gets it done.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-1.png) The Plan agent makes the plan, the Executor agent gets it done. Image by Author But how do you guarantee the quality of what the agents produce in a `plan-execute` setup? That's where a pattern called `reflection` comes in. The reflection pattern adds a reflection agent to the multi-agent workflow. This agent typically runs on a completely different LLM and reviews the output of the `plan` or `executor` agent from a different angle, which raises the overall performance of the multi-agent system. The reflection pattern sees wide use in content creation and deep research scenarios, which proves it works. ![Just adding a reflection step can dramatically improve task quality.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-2.png) Just adding a reflection step can dramatically improve task quality. Image by Author Since the pattern is proven, we can bring the same reflection step into the OpenSpec workflow, targeting the proposal files. I'll introduce a reflection agent that runs on a different LLM from the primary agent, reviewing the proposal files from a different angle. We'll also adjust the OpenSpec workflow so this reflection agent plays an active role in it. And if you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. ### Introducing the reflection agent Adding a new agent in OpenCode is simple. Just drop a new Markdown file into `~/.config/opencode/agents/` with the agent's prompt inside. The core job of this agent is straightforward: review the artifact files that OpenSpec generates from multiple angles, making sure the quality of the requirements input is solid from the start. This agent sits between the `/opsx-propose` and `/opsx-apply` phases. ![Only proposals that pass review can move on to the apply phase.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-3.png) Only proposals that pass review can move on to the apply phase. Image by Author We can have deepseek-v4-pro generate the first draft of this agent's prompt: ```markdown You are an **OpenSpec Change Reviewer** — a critical thinker and auditor focused on substance. Your job is to review every artifact in an OpenSpec change before it moves to implementation, and find the issues that would actually cause implementation failure or rework. ## Core Principle: Distinguish Substantive Defects from Formatting Issues **Substantive defects = issues that cause the implementation to go in the wrong direction, miss critical scenarios, create contradictions, or make acceptance impossible.** **Formatting issues = style or wording differences that don't affect implementation quality.** Your primary job is to find the former. You can mention the latter, but mark them as optional suggestions and put them at the end. ## Your Position You work in the **phase between `/opsx-propose` and `/opsx-apply`**: ``` explore → /opsx-propose → ⬅ you are here (possibly multiple rounds) → /opsx-apply → verify → archive ``` The spec is not yet frozen. Implementation has not started. Your mission: **find the defects that would actually cause rework or incidents before any code gets written**. Catching a spec error takes minutes. Fixing wrong code takes hours. ## Principles - **Constructive and strict.** For every issue, explain not just "what" but "why it would cause rework or an incident." - **Specific, not vague.** Point to exact file locations, requirement names, and task numbers. - **Severity levels.** 🔴 Blocking vs 🟡 Should Fix vs 💡 Suggestion — don't mix them up. - **Context-aware.** Evaluate against the existing system (`openspec/specs/`) rather than in a vacuum. - **Read-only.** Never modify files. You surface problems; OpenSpec executes the fixes. ## Anti-Patterns to Avoid - Rubber-stamping: saying "looks good!" without deep review. - Nitpicking: focusing on formatting while missing architectural flaws. - Jumping to solutions: proposing fixes before the user acknowledges the problem exists. - Ignoring existing specs: reviewing incremental changes without understanding the baseline. - Vague feedback: "this could be better" — say exactly what and why. ``` This is just a partial excerpt of the prompt. The full version is in the source files at the end of the article. To review proposals from a different angle and improve the reflection quality, I recommend using a different LLM for the reflection agent than the one the main agent uses. For example, my primary coding agent uses deepseek-v4-pro, so the reflection agent uses kimi k2.6. ```markdown --- description: OpenSpec Change Reviewer — after propose and before apply, critically reviews all artifact files under the change (proposal/design/specs/tasks) mode: subagent model: kimi-for-coding/k2p6 tools: write: false edit: false bash: false --- ``` ### Locking down the openSpec workflow OpenSpec doesn't actually have a fixed workflow design, users follow a default SDD best practice. So the first step is to lock this workflow down, making OpenCode write an OpenSpec proposal before writing any code. ```markdown --- name: openspec-workflow description: Mandatory prerequisite for ALL OpenSpec operations — load this BEFORE any openspec-* skill. Use when running /opsx-apply, /opsx-propose, /opsx-verify, /opsx-archive, /opsx-explore, /opsx-sync; running `openspec status`, `openspec list`, `openspec instructions`; reading files under `openspec/changes/`; or doing any OpenSpec stage (propose, apply, verify, archive, explore, sync). license: MIT compatibility: Requires openspec CLI. metadata: author: Peng Qian version: "1.0" --- ## OpenSpec Workflow (Mandatory) **All code changes must have a proposal before any code gets written.** ### Process 1. **Explore** - When the user says "think about it," "discuss," or "explore," discuss only — no coding. 2. **Propose** - Create proposal files under `openspec/changes//`. 3. **Apply** - Implement according to the proposal tasks. **No file modifications without a proposal.** 4. **Verify** - Verify after implementation is complete. 5. **Archive** - Archive the change. ### Hard Rules - **Bug fixes don't require editing or creating a proposal.** This hard rule only applies to feature changes or new features. - **No proposal, no change:** If the user asks to modify code, confirm that a matching proposal exists first, or create one. - **No proposal, no edits:** Before editing a file, check that a matching change directory exists under `openspec/changes/`. - **No coding in Explore mode:** When the user is in explore mode, **do not** create proposals, **do not** edit files, **do not** write tests. - **After a change is complete:** Run the verify process to check that the implementation matches the proposal. ### Proposal Creation Requirements Every change must include: - `proposal.md` - reason and scope of the change - `design.md` - design plan - `tasks.md` - specific task list - `.openspec.yaml` - change metadata Additional requirements: - **Task granularity:** Each task in `tasks.md` should take no more than 2 hours. ### Violation Handling Stop immediately and alert the user if any of the following are detected: - Code modification starts without a proposal. - Files are edited in explore mode. - Files outside the current proposal's scope are modified. ``` You can put this workflow into your project's `AGENTS.md` file so OpenCode follows it consistently. Or put it in the global `~/.config/opencode/AGENTS.md` file so you don't have to configure it for every project. A better option is to turn the workflow into a skill, so OpenCode only loads this workflow definition when using OpenSpec. I've packaged the full workflow as the `openspec-workflow` skill, you can grab the source file at the end of the article. ### Inserting the reflection agent into the workflow Once the OpenSpec workflow is locked down as a skill, every SDD coding session will follow the spec-first, code-second process, which means the `/opsx-propose` and `/opsx-apply` phases. As mentioned earlier, the reflection agent sits between these two phases, reviewing the quality of the proposal files before any implementation starts. Following the reflection pattern, this review-and-fix cycle can run for multiple rounds until the proposal artifacts have no serious issues. To prevent an infinite loop, we need a hard cap on the number of review rounds. Let's update the `openspec-workflow` `SKILL.md` file to add the reflection process: ```markdown ### OpenSpec Reflection Process 1. **After each batch of artifacts is created**, the `@openspec-reviewer` agent **must** be called to review the **artifact files in that batch**. 2. The main agent fixes the **current batch of artifact files** based on the feedback from `@openspec-reviewer`. 3. Call `@openspec-reviewer` again to review the **current batch of artifact files**. 4. **Review pass criteria:** 4a. **Single-round pass:** After the current review round, if "### 🔴 Remaining Issues" does not exist or is empty, move to the next batch. 4b. **Fix loop:** If 🔴 issues remain → main agent fixes → next review round → back to 4a. Repeat until passing or 4c triggers. 4c. **Hard cap (MAX_ROUNDS = 5):** If the same batch has gone through 5 review rounds without passing 4a → stop the loop and hand off to a human for a decision. ``` At this point, a typical reflection process for OpenSpec proposals is in place. All you need to do is run `/opsx-propose`, go grab a coffee, and wait for the reflection agent to gradually refine your proposal. ![After the OpenSpec proposal is created, the openspec-reviewer agent will be automatically invoked to begin reviewing and fixing it.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-8.png) After the OpenSpec proposal is created, the `openspec-reviewer` agent will be automatically invoked to begin reviewing and fixing it. Image by Author --- ## Optimizing the Reflection Process After going through a few rounds of "review and fix," you'll notice this reflection process doesn't work as perfectly as you'd expect. It's common to fix issue A only to have issue B pop up. Sometimes all 5 review rounds finish, and not every issue is fully resolved. Especially with powerful but not top-tier models like deepseek-v4\. It's fine. We can optimize the process to fix this. ### Saving the explore discussion to disk OpenSpec has a gap: the `/opsx-explore` phase involves very detailed requirements discussions, but none of that gets saved to a file. Once the user starts a new session or the session hits its context limit, all those discussion details are gone. The review and fix process runs for many rounds and generates a huge amount of context tokens, which almost guarantees the `explore` discussion details get lost. On top of that, OpenCode calls the reflection agent in a separate sub-session to review the proposal artifacts. OpenCode doesn't pass along the background context in detail during that call, so the explore discussion never makes it into the reflection agent's context. So we need to update the `openspec-workflow` skill file to require OpenSpec to generate an `explore-brief.md` document after `/opsx-explore` finishes. This document serves as a persistent checklist baseline for the proposal review process. ![A file-based checklist gives proposals and reviews a solid foundation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-4.png) A file-based checklist gives proposals and reviews a solid foundation. Image by Author We also need to update the reflection agent's prompt to reference this `explore-brief.md` file during the review. Sometimes the requirements are simple and `/opsx-explore` wasn't used, so the primary agent won't generate an `explore-brief.md`. In that case, the reflection agent falls back to the context background passed along when the sub-agent is called. ### Review artifacts one at a time, not all at once In the early stages, I found the biggest reason proposals kept going through so many review rounds was file consistency issues. At first, all proposal artifacts were written at once before any review started. The agent had to review four files in one shot: `proposal.md`, `design.md`, `spec.md`, and `tasks.md`. That meant one review round could surface 8 issues spread across 4 files. Fix file A, and file B becomes inconsistent. Fix file B, and file C needs updating. The next review round then has to check all 4 files again, and the consistency verification cost explodes as N files × N files combinations. This gets especially bad with open-source models like deepseek-v4 or kimi. Fixing A and forgetting B happens all the time. That's the root of the problem. The best fix is to review artifact files in batches. OpenSpec writes `proposal.md`, then immediately enters the review phase. Once it passes, `proposal.md` is frozen and can't be changed. Then OpenSpec generates `design.md` and enters review. Once that passes, `design.md` is frozen, then `spec.md` is generated and reviewed, and so on. This way, each subsequent artifact's review naturally includes a consistency check against the already-frozen earlier artifacts. Why does this guarantee consistency? Let me draw a diagram: ![Each new batch of artifacts uses the previous frozen batch as its reference.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-5.png) Each new batch of artifacts uses the previous frozen batch as its reference. Image by Author The consistency guarantee comes from two mechanisms: 1. The freeze mechanism. Artifacts that pass review can't be freely modified anymore. This eliminates the back-and-forth of fixing A, then fixing B, then having to fix A again. It cuts down on a lot of repeated review rounds. 2. One-way dependency. proposal → design → specs → tasks is a strict DAG. When reviewing design, you only need to check its consistency with the proposal, not with specs (since specs haven't been written yet). Each review round only compares one pair (new artifact vs. frozen artifact), not N × N. Generating and reviewing artifacts in batches may not reduce the total number of rounds compared to generating everything at once, but each round only touches 1 or 2 files. Fixes don't accidentally break other files, and you don't have to worry about the LLM fixing A and forgetting B. The final review process looks like this: ![The final proposal review flow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-6.png) The final proposal review flow. Image by Author ### Log every review round Same problem again: OpenCode calls the reflection agent in a sub-session to review proposals. Not only does the sub-session lose its input context (already solved with `explore-brief.md`), it also loses the history of previous review conclusions. The primary agent starts fixing the current round right after getting the reflection agent's feedback, without saving those results to the session context. As the review and fix context grows, the LLM's attention inevitably drifts. The best fix is to have the primary agent log each review's issues and fix results. That way, in later review rounds, the LLM can get a full picture of what happened historically, what decisions were made, and how they affect the current round. And if the review hits the max rounds and needs human intervention, we have a much better reference to work from. The implementation is simple: after each round of fixes, append the current round's review results to a `review-log.md` file. ![The main agent should log the results of every proposal review round.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/05/image-7.png) The main agent should log the results of every proposal review round. Image by Author --- ## Conclusion Let's recap everything we did to improve OpenSpec proposal artifact quality through the reflection mechanism: 1. We set up a reflection agent. Its job is to fully review the proposal artifacts that OpenSpec produces before the `/opsx-apply` phase starts, making sure everything is solid before moving forward. 2. We locked down the OpenSpec workflow as an `openspec-workflow` skill. This skill requires OpenCode to strictly follow the spec-first, code-second process. After the proposal is generated, the reflection agent is called to review it. 3. We patched a gap in OpenSpec by requiring it to record each `/opsx-explore` requirements discussion in an `explore-brief.md` file. This prevents the LLM's context window from losing important details during long review processes. Each review and fix result is also recorded incrementally in a `review-log.md` file, for the same reason. 4. We changed how OpenSpec generates artifact files: from generating all files at once to generating one file, reviewing it, and freezing it before moving to the next. This eliminates the root cause of fixing A, breaking B, fixing B, then having to fix A again. With these reflection harness optimizations in place, even using deepseek-v4-pro, code bugs are rare now. The model covers all kinds of edge cases thoroughly, code reviews almost always pass on the first try, and messy code is much less common. That said, adding this reflection harness does make the OpenSpec proposal generation process much longer. It takes more waiting time and extra tokens. But getting DeepSeek-v4 to match or even beat Claude Opus in code generation quality makes it worth it overall. Beyond the reflection mechanism, there's still a lot of room to improve the OpenSpec coding workflow. I'm designing a self-evolving plugin system that lets agents learn from each review and fix cycle, so OpenCode can avoid making the same mistakes in future proposal generation. I look forward to walking you through that in the next article. Thanks for reading and subscribing. If anything in the article wasn't clear, leave a comment and I'll get back to you. Feel free to share this article with your friends, it might help someone out. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further reading Give my programming workflow built with OpenCode, OMO-Slim, and OpenSpec a try? [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-5d6af005-6ad4-4624-bb8f-506daa74410b.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-a6ae52b4-54e9-48ce-a04f-24458b64d76b.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) Can't your DeepSeek-V4 and GLM-5.2 agents read images yet? Give my method a try: [DeepSeek-V4 Can’t Read Images? I Made It ReadDon’t wait for a multimodal model, you can use it now![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-31dee510-7149-4c48-904a-8fbc8a520bd3.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-72ad4786-cf6a-4a06-9060-8be468016065.webp)](https://www.dataleadsfuture.com/deepseek-v4-cant-read-images-i-made-it-read/) Better code quality, but spend less money. [How I Cut Kimi K3 Costs in OpenCodeBetter code quality, but spend less money![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-2ebe8c1c-7051-4fc5-960a-0741067de3b9.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover_tiny-0af579bb-8e82-4d2f-bcbb-d5a38796acd7.webp)](https://www.dataleadsfuture.com/how-i-cut-kimi-k3-costs-in-opencode/) --- ## Source Code Here's the source code for this article. Sign up now to get it for free. [Grab the Source Code ](#/portal/signup) ### How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding Environment URL: https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/ Last updated: 2026-07-31T08:05:53.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* This article was last updated on July 23, 2026. ## Introduction I have never used Claude Code. The reason is simple. Claude Code is too expensive. Even with a subscription, the cost-to-value ratio does not work for my research. So I have been building my own AI coding environment using OpenCode as the foundation, combined with `Oh-My-OpenCode-Slim` (multi-agent orchestration) and `OpenSpec (SDD)`. My take is this: if you understand what you want to build, and you know how to use coding tools properly, especially with well-written Spec files as constraints, frontier open-source models like `Qwen3.6-Plus`, `Kimi-k2.5`, and `GLM-5` can handle your daily coding tasks just fine. There is another huge advantage to open-source software: community power. The community can tune system prompts and model parameters to fit different models well, and get the most out of open-source models. In this article, I want to share what I have learned from using OpenCode and its surrounding tools. **I will skip the generic tutorials you find everywhere online and focus only on the details I think actually matter.** I hope this helps you make better choices. --- ## Tool Installation and Environment Setup I will cover my experience in two parts: installing and configuring OpenCode and its related plugins, and my AI coding workflow. Let's start with tool installation and environment setup, beginning with OpenCode itself. ### Installing OpenCode Unlike most coding agents that only offer a TUI-based command-line tool, OpenCode also comes with a desktop app with a graphical interface. I use the desktop app for my daily coding work. It is clearly much more efficient than the TUI version. That said, you still need to install the command-line program first. From my testing, some plugins need a command-line environment to check whether OpenCode is installed on your machine during project initialization. #### Why does the installation fail OpenCode relies on `bun` as its only runtime. But in a lot of workplace environments, restrictions can prevent `bun` from being installed automatically. In that case, you just need to install `bun` manually: ```bash npm install -g bun ``` ### Configuring OpenCode After installing OpenCode Desktop, open the app. Once you select your project directory, you will land on the main OpenCode interface. The features are fairly intuitive, so I will not walk through each one. But before you type your first `Hello World`, you should check your terminal configuration first. #### Configuring the terminal I use Windows 11\. On Windows, OpenCode Desktop defaults to PowerShell as its terminal. Many companies, though, do not allow PowerShell. If you are in a non-English locale, OpenCode may run into character encoding issues when running shell commands through PowerShell, causing those commands to fail. In that case, you need to change your default terminal. OpenCode uses the `SHELL` environment variable to determine which terminal to use. You can configure Windows Command Prompt (`cmd.exe`), WSL, or `git bash`. Personally, I prefer `cmd.exe` because I had already installed a lot of CLI tools before setting up OpenCode. Using `cmd.exe` directly saves me from reinstalling everything. ```shell SET SHELL="%windir%\system32\cmd.exe" ``` #### Configuring model providers Next, let's talk about how to configure model providers. Open the settings window and select "Providers". You will see a list of the most popular model providers and API relays. If you want to use open-source models, though, they probably will not be on that list. At that point, you might click "Custom provider" and manually fill in the `model id`, `base url`, `api key`, and so on. The problem is that OpenCode then has no idea about your model's context window size or pricing, which causes features like automatic context compression to stop working correctly. The right approach is to click the "Show more providers" link at the bottom, find the provider you want to add, and enter that provider's `api key`. ![Click the show more providers link and pick a provider for your open-source model.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image.png) Click the show more providers link and pick a provider for your open-source model. Image by Author Once configured, all models from that provider will appear in the model list. These models come with metadata like context size and pricing, so context management plugins can work correctly. ![By setting up your provider correctly, you'll get all kinds of metadata about your models. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-1.png) By setting up your provider correctly, you'll get all kinds of metadata about your models. Image by Author The downside is that you cannot directly see your provider ID this way, which makes it tricky to configure `Oh-My-OpenCode-Slim` later. No worries. OpenCode already saved your provider configuration when you selected your provider. You can find it at `~/.local/share/opencode/auth.json`. Your provider ID and API key are both there. ![You can find your provider ID in the auth.json file.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-2.png) You can find your provider ID in the `auth.json` file. Image by Author #### Enable workspaces The biggest difference between AI coding and traditional coding is that while you wait for the AI to work, you can actually work on another requirement at the same time. If you use git for version control, you would normally need to create a separate directory and check out a new branch. Or you can use git's worktree feature to create a new worktree on top of your current branch. When you are doing parallel development, using worktrees is much more convenient than creating new branches. Compared to OpenCode CLI, the desktop app has a clear advantage here: it natively supports the worktree feature. In OpenCode Desktop, this feature is called "workspace". The way to open a workspace is a bit hidden. Right-click the project icon in the top-left corner of the window, then select "Enable Workspace" from the menu. From there, you can create multiple workspaces in the conversation list and work on them simultaneously. The corresponding branches and code directories will be created automatically. ![Right-click on your project to enable the workspace.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-3.png) Right-click on your project to enable the workspace. Image by Author When the coding work in a workspace is done, you can ask the AI to submit a PR for the current code, then close the workspace. The branch and code directory that were created for it will be cleaned up automatically. Very convenient. #### Choosing the right agent If you want to use OpenCode without any extra plugins, pay attention to how you use agents. OpenCode has two types of agents: primary agents, which you choose yourself, and sub-agents, which the primary agent calls on its own when needed. Without any plugins installed, OpenCode only provides two primary agents: `Build` and `Plan`: The `Build` agent has full tool access and is the default choice. The `Plan` agent has no editing permissions. Its job is to ask you clarifying questions when you describe a requirement, and eventually produce an execution plan. When you first try this, you might go straight to the `Build` agent. But for complex tasks, `Build` tends to just start coding based on its own interpretation. That is like looking through a straw. It fixes things locally without thinking through the overall architecture and design patterns. The right approach is to start every new requirement with the `Plan` agent for requirement clarification, and get a solid execution plan first. Only then should you hand things off to `Build` to start development. But even that is not enough. Model context is limited. As coding progresses, the execution plan from earlier in the conversation can get pushed out of the context window. A better approach is to ask `Build` to save the execution plan as a Markdown file before starting to code. Review that file, confirm everything looks good, then start a fresh session and have `Build` load the execution plan document back in before executing. ![The Plan agent generates a tasks file, and the Build agent executes them.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-4.png) The Plan agent generates a tasks file, and the Build agent executes them. Image by Author Once you start working this way, you will feel how much value comes from planning before executing. It also sets you up well for SDD coding later, which I will cover when we get to OpenSpec. If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. #### Always start a new session I just mentioned that after forming a development plan, you should create a new session before continuing with coding. Why? Because anyone familiar with LLMs knows that even though context windows are long today, and OpenCode does offer context compression, context rot is still a real problem. LLMs pay more attention to the beginning and end of the context window, and less to the middle. I call this positional bias. So to make sure the LLM follows instructions accurately based on the conversation context, especially after forming an execution plan where you need precise execution, start a new session after each major milestone. Do not keep working in the same session forever. #### Enable the Plan Agent's Planning Workflow Mode By default, the `Plan` agent runs in a pure read-only analysis mode. In this mode, it just has a back-and-forth with you about requirements and how to implement them through heuristic reasoning, but it can't edit any files. This creates a real problem: everything discussed in `Plan` mode gets scattered across the current session's chat history. Over time, the context starts to rot, and once you switch to a new session, all that discussion is just gone. As spec-driven development has proven its value, starting from v1.15.0, OpenCode introduced an experimental planning workflow mode for the `Plan` agent. This workflow follows 5 phases: 1. Explore. When a new task comes in, the `Plan` agent spins up to 3 `explore` agents in parallel to analyze the codebase. (I'd recommend configuring the `explore` agents with the `deepseek-v4-flash` model here, otherwise the token costs can get pretty painful.) 2. Design. Once the codebase is understood, `Plan` calls a general sub-agent to design the development approach. 3. Review. `Plan` agent checks the development plan against your original task intent to make sure it actually holds up. 4. Write the plan file. The final plan gets written to a `.md` file. This is the only file type the `Plan` agent is allowed to write to. 5. Exit Plan mode. It calls the `plan_exit` tool and lets you know it's time to leave `Plan` mode and start writing actual code. ![Planning workflow mode sequence diagram.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/opencode-plan_workflow.drawio.png) Planning workflow mode sequence diagram. Image by Author Where `Plan` agent stores its plan files: - When a `git` repo is present: `/.opencode/plans/-.md` - When there's no `git` repo: `~/.local/share/opencode/plans/-.md` To enable the planning workflow mode for the `Plan` agent, set the following environment variable: ```bash SET OPENCODE_EXPERIMENTAL_PLAN_MODE=true ``` If you want a quick taste of OpenCode's native SDD development experience, just flip this environment variable on. No extra plugins needed. But if you're looking to use SDD for production-level AI Coding, I'd suggest checking out the OpenSpec configuration covered later in this article, paired with the [Reflection SDD workflow](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) I put together, to get truly top-tier development quality. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. #### Do not forget to create`AGENTS.md` The `AGENTS.md` file is called "rules" in OpenCode. You can create it automatically with the `/init` command. It tells the LLM what rules to follow during coding. You may ask: if this file is created by the LLM, does that not mean the LLM already knows all these rules internally? Is saving them to a file redundant? Not at all. `AGENTS.md` is a file written specifically for the LLM to read. In my view, it serves three important purposes: First, `AGENTS.md` acts as long-term memory for the project. It locks in facts and choices. For example, after asking the LLM to set up the project structure or create a new module, run `/init` once. The project architecture gets locked into `AGENTS.md`. Without this, the LLM will scan the entire project from scratch every time you start a new session, wasting a huge number of tokens. Another example: if you use `uv` to manage your project and use `uv sync --prerelease=allow` to sync prerelease dependencies, write that clearly in `AGENTS.md`. This prevents the LLM from making errors with dependency management. Second, `AGENTS.md` narrows the probability distribution and reduces hallucinations. LLMs are probability models. When facing a question, an LLM generates several possible answers with associated probabilities, then picks one based on the `temperature` parameter. For example, when a method parameter can accept multiple types, the LLM might consider these options: 1. Use `Optional[int]` (40% chance) 2. Use `int | None` (40% chance) 3. Use no type annotation at all (20% chance) At that point, the LLM will randomly pick `Optional[int]` or `int | None`. But once you explicitly require `str | None` syntax in `AGENTS.md`, the probability distribution shifts to: 1. Use `int | None` (100% chance) 2. Use `Optional[int]` (0% chance) At this point, the LLM will just go ahead and pick `int | None` as the final answer. Third, `AGENTS.md` serves a harness engineering purpose. You can give the LLM direct instructions through `AGENTS.md` that it must follow. For example, you can tell the LLM to communicate with you in Chinese, or require that during the spec-driven process, it cannot create new proposals without your approval. #### Improve the success rate of Skills loading By now, most people in the AI coding space have heard of Skills. But many find that Skills do not load reliably under normal conditions. There are two reasons for this. On one hand, the `description` in a Skill's front matter is often unclear. We need to write the front matter carefully, especially the `description` part. It should clearly explain when the Skill applies and what it provides, so the LLM can load the right Skill for the situation. On the other hand, for common coding scenarios, LLMs have learned so much during pretraining that they do not feel the need to load a Skill for extra guidance. For this, there is a simple and proven fix. Just add this line to `AGENTS.md`: > Prioritize retrieval-led reasoning over pretrained-knowledge-led reasoning. That is all. After receiving this instruction, the LLM will load the relevant Skill for a given coding scenario instead of falling back on its internal pretrained knowledge. From my testing, the Skill loading success rate jumps from around 60% to 90%. On top of that, with this instruction in place, the LLM will use tools like `glob/grep` to check existing code structure before starting, and will use `websearch` more often to look things up online. It follows a "search first, verify after" approach instead of relying on its own instincts. Definitely worth trying. ### Installing and Configuring Oh-My-OpenCode-Slim Now that we have covered the key details of OpenCode itself, let's move on to `Oh-My-OpenCode-Slim`. `Oh-My-OpenCode-Slim` is a multi-agent orchestration plugin. It provides six types of agents, each focused on a specific coding scenario. | Agent | Scenario | | ------------ | -------------------------------------------------------------------- | | Orchestrator | Coordinates other agents to complete tasks together. | | Explorer | Explores the project codebase to understand what has been done. | | Oracle | Fixes complex bugs or provides architecture-level advice. | | Librarian | Searches GitHub repositories and framework API documentation online. | | Designer | UI design and user experience. | | Fixer | Handles coding work in parallel. | You might ask: why `Oh-My-OpenCode-Slim` instead of `Oh-My-OpenCode`? Because at this stage of AI coding, multi-agent orchestration workflows have become too heavy and less important than they used to be. A lightweight multi-agent plugin like `Oh-My-OpenCode-Slim` fits much better here. Each agent's `system prompt` is concise. The main value comes from the efficiency of parallel execution across agents, and the context isolation that sub-agents provide, which helps the primary agent stay focused. It does not add much to your token costs. #### Installing the plugin Let's start with what to watch out for when installing `Oh-My-OpenCode-Slim`. Even if you use OpenCode Desktop, you need to install OpenCode CLI first. Otherwise, the installation of `Oh-My-OpenCode-Slim` will fail. This is because `Oh-My-OpenCode-Slim` checks for the OpenCode CLI to determine whether OpenCode is installed on your machine. Due to a series of recent npm package poisoning incidents, OpenCode removed the feature that auto-installs plugins through `opencode.json` configuration. You now need to install the latest version of `Oh-My-OpenCode-Slim` using this command: ```shell bunx oh-my-opencode-slim@latest install --no-tui --tmux=yes --skills=yes ``` This command installs the `Oh-My-OpenCode-Slim` agents and also installs two Skills: `simplify` and `agent-browser`. The former simplifies and cleans up unnecessary code. The latter helps agents inspect the frontend page styling. You will also notice that in the `opencode.json` config file, the `Oh-My-OpenCode-Slim` plugin points to a local entry file instead of an npm package. So for future updates, you will need to update it manually. #### Configuring the plugin Now let's talk about provider config. The official documentation covers this in reasonable detail, so just follow the configuration for your model. If you configured your model through the provider list, you cannot directly see the Provider ID. No worries. As I mentioned earlier, `~/.local/share/opencode/auth.json` stores all the providers you have configured, including the provider IDs. ```json { "preset": "kimi", "presets": { "kimi": { "orchestrator": {"model": "kimi-for-coding/k2p6", "skills": ["*"], "mcps": ["*"]}, "oracle": {"model": "kimi-for-coding/k2p6", "variant": "high", "skills": ["simplify"], "mcps": []}, "librarian": {"model": "kimi-for-coding/k2p6", "variant": "low", "skills": [], "mcps": ["websearch","context7","grep_app"]}, "explorer": {"model": "kimi-for-coding/k2p6", "variant": "low", "skills": [], "mcps": []}, "designer": {"model": "kimi-for-coding/k2p6", "variant": "medium", "skills": ["agent-browser"], "mcps": []}, "fixer": { "model": "kimi-for-coding/k2p6", "variant": "low", "skills": [], "mcps": []} } } } ``` After configuring, remember to run `opencode auth login`. Even if you already filled in your `API KEY` through the provider list earlier, the agents will not connect properly unless you run this CLI command. #### Trying out the Council feature `Oh-My-OpenCode-Slim` introduced a `Council` agent in its latest version. More accurately, `Council` is a collection of agents. It sends the same task to several different agents that make up the `Council`, calculates a confidence interval from their responses, and synthesizes a final answer. It is a bit like ensemble learning in machine learning. "Two heads are better than one," and by using this parliamentary discussion approach, even a few different open-source models working together can match or even exceed the capability of `Opus 4.6`. This makes it a great fit for solving complex architecture design problems or tracking down tricky bugs. `Council` configuration is straightforward. Just follow the official example. For best results, pick a few different agents as `Council` members to make the most of the ensemble learning approach: ```json { "council": { "master": { "model": "alibaba-cn/glm-5" }, "presets": { "default": { "alpha": { "model": "kimi-for-coding/k2p6" }, "beta": { "model": "alibaba-cn/qwen3.6-plus" }, "gamma": { "model": "alibaba-cn/MiniMax-M2.5" } } } } } ``` OpenCode's UI was not designed for this parliamentary-style interaction, so the `Council` discussion process is not very transparent. If something goes wrong in the configuration, you cannot tell from the interface. We need a way to debug `Council`. The method is simple. Since `Council` itself is a primary agent you can select from the agent list, just select `Council`, then type "test Council connectivity" in the chat. OpenCode will send a test task to each `Council` member and list the results from each one in detail. Very handy. ![Once you've got Council set up, you can test the connectivity of each sub-agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-5.png) Once you've got Council set up, you can test the connectivity of each sub-agent. Image by Author You can use this same approach to test other OpenCode features. For example, type "test `context7` availability" or "test `Oracle` agent connectivity." #### Why isn't your LSP tool working anymore? In the earlier versions, the omo-slim agent used to run a tool called `lsp_diagnostics` after writing or editing code to check for syntax issues, making sure there were no unused imports or dead code and stuff like that. I really liked that tool, it was pretty handy. But after the latest omo-slim plugin update, you might notice the agent no longer calls `lsp_diagnostics` to check the code. That's because the old omo-slim used its own built-in LSP tool, while the new omo-slim is going to use the LSP tool provided by opencode. However, opencode has the LSP tool disabled by default, since the author thinks it uses up too much memory and doesn't really pay off in most situations. If you still want the agent to use `lsp_diagnostics`, you can manually add the `"lsp": true` setting in the opencode.json config file, and that'll turn it back on. Also, opencode offers other LSP tools like `goToDefinition`, `findReferences`, `hover`, `workspaceSymbol`, `goToImplementation`, and so on, which are useful for code searching and debugging. You can enable them by setting the environment variable `OPENCODE_EXPERIMENTAL_LSP_TOOL = true` and give them a try. But even with those tools enabled, the agent will still prefer using `grep` for text based code searches. That's because the LLM was trained on data that heavily uses `grep` during the pretraining phase. So from my own testing, turning on the LSP tools didn't make a big difference. ### OpenSpec and SDD Finally, let's look at another key piece of the AI coding puzzle: the `OpenSpec` plugin and the SDD coding workflow. Why do we need SDD? The reason is pretty simple. A lot of articles that benchmark model coding ability use prompts like "build a 3D web demo from this one-sentence description," then compare results and conclude that open-source LLMs still have a long way to go. But is that really true? Think back to how software development worked before AI coding. Before writing a single line of code, did we not always ask the business side for detailed requirements, write design documents, plan the coding schedule and testing plan, and only then start writing code? So why would anyone think you can skip the documentation phase when asking an AI to write code? ![Think things through with the docs first, then roll up your sleeves and get to work.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-6.png) Think things through with the docs first, then roll up your sleeves and get to work. Image by Author The ability to write coding specs that make an LLM work with precision is one of the golden standards for evaluating a programmer's skills in the AI era. No need to stress about it, though. There are plenty of AI tools today that help you do SDD development. `Superpowers`, `BMAD`, and `OpenSpec` are among the best. Compared to `Superpowers` or `BMAD`, `OpenSpec` is much lighter. For most everyday projects, `OpenSpec` is more than enough to get the job done. Let's look at how to install and configure it. Even though we use OpenCode and `OpenSpec` together, `OpenSpec` is not a plugin for OpenCode. It is a completely independent tool. #### Configuring the workflow The default `OpenSpec` workflow has three phases: `opsx-propose`, `opsx-apply`, and `opsx-archive`. These correspond to creating a proposal, implementing the proposal, and archiving the proposal. Just like I recommended using the `Plan` agent to plan things out before implementing any changes, I also want you to have a thorough discussion with the AI about your requirements before creating a proposal. Make sure both you and the AI clearly understand what you want to build. Since `OpenSpec` does not have the concept of agents, the responsibility for requirement clarification falls on a command called `opsx-explorer`. Compared to the `Plan` agent, the `opsx-explorer` command uses a Skill to strengthen the ability to guide users through their requirements in an exploratory way, helping them think things through more completely. `opsx-explorer` is also great at using charts to compare the pros and cons of different implementation options, making decisions easier for the user. `opsx-explorer` is not included in the default workflow, so you need to configure `OpenSpec` to add it. After installing `OpenSpec`, run `openspec config profile` in the command line, then select `Workflows only` in the configuration screen, as shown below: ![You just need to set up the Workflow, that's it.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-7.png) You just need to set up the Workflow, that's it. Image by Author Use the spacebar to select `Explorer ideas` and `Verify change`. These correspond to `opsx-explorer` and `opsx-verify`. Then press `Enter` to confirm. ![Just use the Space key to select the workflow nodes you like.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-8.png) Just use the Space key to select the workflow nodes you like. Image by Author After configuration, go to your project root directory and run `openspec init` to set up a new `OpenSpec` environment, or run `openspec update` to update an existing `OpenSpec` configuration. I will walk through the full `OpenSpec` development workflow in the next section. #### Configuring multi-language support If you try creating rule files with `OpenSpec`, you will notice they are written in English by default. That is inconvenient for non-English users. So we need to configure `OpenSpec` to write files in a specific language. A simple approach is to add a language instruction directly in `AGENTS.md`, but that causes a problem. `spec.md` files use a lot of English keywords. If you tell `OpenSpec` to use another language, those keywords get translated too, and the LLM can no longer follow the spec constraints correctly. The right approach is to configure the language in the `config.yaml` file inside the `OpenSpec` directory: ```YAML schema: spec-driven context: | Language: Chinese Write in Chinese, but: - Keep technical terms like "API", "REST", "GraphQL" in English - Code examples and file paths remain in English Tech stack: Python 3.13 Rules: proposal: - Keep proposals under 1500 words - Always include a "Non-goals" section tasks: - Break tasks into chunks of max 2 hours ``` Besides language, you can also configure other settings here, such as your tech stack and specific requirements for `proposal.md` and `tasks.md` files. #### Choosing the right agent Since `OpenSpec` needs to write spec-related files, the agent must have file editing permissions. If you use the original `OpenCode` tools, you should go with the `Build` agent, since the `Plan` agent has no editing capability. In this tutorial, though, we have integrated the multi-agent orchestration from `Oh-My-OpenCode-Slim`. So you should use the `Orchestrator` agent. Beyond basic editing, the `Orchestrator` can call `Explorer` and `Librarian` agents during the planning phase to analyze project structure and search API documentation. During execution, `Orchestrator` can also call multiple `Fixer` agents to work in parallel, which speeds up development significantly. From now on, just default to the `Orchestrator` agent. It will call other agents when it needs to. --- ## Daily OpenCode Development Workflow Now that we have covered installation and configuration, let me walk you through my daily `OpenCode` workflow using `OpenCode`, `Oh-My-OpenCode-Slim`, and `OpenSpec` together. ### Project Initialization Unlike traditional coding workflows, the development process in the AI era should center on building for AI. All configuration and documentation should be written in a way that AI can understand and follow. The AI then translates human ideas into code. Let me walk you through this from project initialization. After creating the project, first run `git init` in the project root directory. If the project has a remote repository, remember to link it with `git remote add`. At that point, you can start using OpenCode's code review and workspace (worktree) features. Next, if your project is a Python project, use `uv init` to initialize dependencies and project structure. If needed, also add lint and code review tools with `uv add --dev ruff` and `uv add --dev mypy`. This gives the LLM a clear anchor for understanding the project. `pyproject.toml` is the standardized build protocol introduced by `PEP 518`, and most build tools use this file. Without running `uv init` upfront to establish a basic project structure, the LLM will randomly pick `pip`, `poetry`, or something else. Fixing that later is a pain. Next, initialize the `OpenSpec` plugin. Still in the project root, run `openspec init`. In the command-line interface that appears, select the `OpenCode` agent and press Enter. Finally, go into the `openspec` directory and edit `config.yaml` to add your rules, as described earlier. Once everything is configured, run the `/init` command and ask OpenCode to generate the `AGENTS.md` file to lock in all these configuration choices. ### The Development Workflow With the project initialized, you can now move into the actual development flow. As I described earlier, you should have already configured the `OpenSpec` workflow using `openspec config profile` and selected `opsx-explorer` and `opsx-verify`. So the best way to start is to open a new session and begin with `/opsx-explorer`. No matter how clear your idea feels, use `/opsx-explorer` to flesh it out, fill in the details, and help the AI clear up any questions. I know you want to just tell the AI to build a Facebook. But I recommend starting small, one small requirement at a time. For example: "I want to build a budgeting app. For step one, let's start by setting up the project structure and adding dependencies." This keeps things focused. The AI completes a small amount of work at a time, which avoids context limitations, causing the AI to go off track after a long run. Once the requirement discussion is mostly done, run `/opsx-propose` to create a proposal. This locks down the discussion and development plan as spec documents. `OpenSpec` will break things down into one or more changes based on complexity, then generate `proposal.md`, `design.md`, and `tasks.md`. Review these three files carefully. The LLM-generated tasks might have gaps or misunderstandings. If you find errors, you have two options: - Talk to the LLM, add details, and ask `OpenSpec` to regenerate the documents. - Or edit the rule files directly, then ask `OpenSpec` to regenerate the documents. Once the documents look good, I strongly recommend starting a new session with a fresh context. Then run `/opsx-apply` to enter the implementation phase. The agent will load the spec documents from the previous session and start coding. With `Oh-My-OpenCode-Slim`'s multi-agent system and various Skills helping out, this process mostly runs on autopilot. Go grab a coffee or read another one of my articles while you wait. After coding is done, `OpenSpec` will prompt you to run `/opsx-archive` to archive the proposal. But before archiving, I recommend switching to a different model. For example, if you were using `kimi k2.5`, switch to `GLM-5`, then run `/opsx-verify` to verify that all tasks were completed. You can also run `/review` or `/simplify` to review and clean up the code and make sure the project quality is solid. ![A typical SSD development cycle.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/04/image-9.png) A typical SSD development cycle. Image by Author Finally, run `/opsx-archive` to archive the changes. One standard AI development cycle is now complete. When the next new requirement arrives, start fresh with a new session and `/opsx-explorer`. --- ## Conclusion In this tutorial, I covered how to use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to build your own AI coding workflow. With this open-source combination, paired with the latest open-source models, you can get coding performance that rivals top-tier models, at a much better price. This tutorial draws from my real-world experience. It focuses on the practical details of AI coding with OpenCode and walks through each step of the SDD-based development cycle. By the end of this tutorial, you should have a solid edge in the AI coding space. Due to space and my own limitations, I could not cover every detail. I am still working on improving this article. If anything is unclear as you read through it, leave me a comment, and I will get back to you as soon as I can. Thanks for subscribing. Feel free to share this article with your friends. You never know who it might help. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further reading By adding a reflective agent to the OpenSpec workflow, I managed to get DeepSeek-V4-Pro to perform at the level of Opus: [Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec WorkflowStop letting bad spec files tank your code quality![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-8e64abbc-0abf-4933-9510-28270569186f.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/9f04a900-aa9a-4fd2-bd70-6f3fe3783129-4b663419-d8b8-475a-994c-ee00fb661600.webp)](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) Can't your DeepSeek-V4 and GLM-5.2 agents read images yet? Give my method a try: [DeepSeek-V4 Can’t Read Images? I Made It ReadDon’t wait for a multimodal model, you can use it now![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-ab31ee0f-a18e-47c3-a30a-dbbb78204849.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-c75bbbb4-131d-440a-94e6-2f766a866d4f.webp)](https://www.dataleadsfuture.com/deepseek-v4-cant-read-images-i-made-it-read/) The concept of Loop Engineering has been getting a lot of buzz lately, so I decided to give it a shot in OpenCode. The results were surprisingly good: [No Plugins Needed, I Built a Fully Automated Coding Loop in OpenCodeUsing DeepSeek-V4 for low-cost Loop Engineering![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-403fd322-1bcf-485f-b099-965b43d45c2d.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-3-comprass-32a4ac57-5d2c-475a-a8f8-6a0a837dae5b.webp)](https://www.dataleadsfuture.com/no-plugins-needed-i-built-a-fully-automated-coding-loop-in-opencode/) Better code quality, but spend less money. [How I Cut Kimi K3 Costs in OpenCodeBetter code quality, but spend less money![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-dba613f8-0702-422d-a175-81be6d1e073a.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover_tiny-a88bc7d5-d4cd-4249-9378-7a6eaf0073c6.webp)](https://www.dataleadsfuture.com/how-i-cut-kimi-k3-costs-in-opencode/) ### How to Use Agent Skills in Enterprise LLM Agent Systems URL: https://www.dataleadsfuture.com/how-to-use-agent-skills-in-enterprise-llm-agent-systems/ Last updated: 2026-07-27T01:41:24.000Z ## Introduction Enterprise-grade agentic systems have fallen way behind the desktop agent apps that everyone's been buzzing about lately. After spending the better part of a year building enterprise agent applications, I came to one conclusion: **if your agent system can't plug into your company's existing business processes, it won't bring real value to your organization.** Desktop systems like OpenClaw and Claude Cowork solved this problem. They don't change their agent setup at all. Instead, they use Agent Skills to capture human business processes, then share those skills between desktop agent systems through the file system. That's how they tackle one business problem after another. But enterprise users write their skills through a web interface and save them to a database. There's a good chance the process involves complex approval and security audit steps, too. So how does your agent load these skills in real time without any downtime? The latest version of Microsoft Agent Framework finally makes this possible with its Agent Skills feature. ### TL;DR With Agent Skills in Microsoft Agent Framework, enterprise agent systems can load user-defined business process skills from a database in real time, and run the scripts and generated code that come with those skills safely inside containers. Your agent system stays secure and stable, while gaining the same flexible business process orchestration that desktop agents enjoy. All the source code in this tutorial is available at the end of the article. --- ## Before We Start ### Install the latest Microsoft Agent Framework To use Agent Skills, install the latest version of Microsoft Agent Framework: ```shell pip install agent-framework --pre ``` Or, like me, you can pin the version of `agent-framework` in your `pyproject.toml`: ```letax dependencies = [ "agent-framework>=1.0.0rc4", "agent-framework-ag-ui>=1.0.0b260311", ] ``` Then tell `uv` to allow prerelease versions: ```shell uv sync --prerelease=allow ``` ### Install Tavily Agent Skills My end goal is to show you how to share and load Agent Skills between agents deployed across distributed nodes. But I think we should start simple. First, let me show you how to load and use skills from the community. Let's start with Tavily Agent Skills. We'll only load the `tavily-best-practices` skill. It guides my agent on how to generate Tavily-based search code based on the task at hand, instead of calling a hardcoded function: ```shell npx skills add tavily-ai/skills ``` Don't worry. After the initial demo, I'll walk you through how to load skills from a database in real time. --- ## How to Load Agent Skills from Disk Let's start with the most basic approach. In Microsoft Agent Framework, context operations are handled by a base class called `ContextProvider`. The latest version of MAF ships a `SkillsProvider` class. Use it directly and pass the location of your skills through the `skill_paths` attribute, and you're done. `skill_paths` doesn't require a default directory like `.claude/skills`, and you can pass in multiple paths. ```python skills_provider = SkillsProvider( skill_paths=get_current_directory() / ".agents/skills", ) ``` Next, create your agent and pass the `skills_provider` instance through `context_providers`. ```python skills_agent = chat_client.as_agent( name="SkillsAssistant", instructions="You're a helpful assistant, and you'll respond to user requests according to your skills.", context_providers=[skills_provider], tools=[code_tool], ) ``` To run the Python code the agent writes based on the Tavily skill instructions, you need to pass a `code_interpreter` tool to the agent. Let the code run inside a container environment. I'll cover that in detail later. Write a `main` method to test the agent: ```python async def main(): async with code_executor: session = agent.create_session() result = await skills_agent.run( "Check how gold ETFs performed in February 2026 and give some investment advice.", session=session ) print(result) ``` Microsoft Agent Framework provides an OpenTelemetry-based telemetry tool. I hooked it up to MLflow. Let's run the agent once and see what happens: ![Through MLflow, you can see that the agent successfully loaded and executed the Agent Skills.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image.png) Through MLflow, you can see that the agent successfully loaded and executed the Agent Skills. Image by Author You can see that once the agent decided it needed Tavily to search, it loaded the full `SKILL.md` document, wrote Tavily search code following the instructions, then sent it to the code interpreter for execution. Exactly what we expected. You can learn how to use MLFlow in this article: [Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracing TutorialEnhance your multi-agent application’s observability, explainability and Traceability![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-66.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Tracing-with-MLflow-3.x-1-2.webp)](https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/) --- ## How Agent Skills Work Now let's talk about how to get the most out of Agent Skills in enterprise systems. That means loading external skills in real time, containerizing the code interpreter, and managing context more carefully. But before we go there, let's dig into how Agent Skills actually work inside MAF, so the rest of this tutorial makes more sense. As I mentioned, `SkillsProvider` extends `BaseContextProvider`, which means it works by operating on the agent's context. When you initialize `SkillsProvider`, you pass one or more search paths to the `skill_paths` attribute. Take the `.agents/skills` directory as an example. On startup, `SkillsProvider` recursively searches this directory and finds every subdirectory that contains a `SKILL.md` file. Then it extracts the `name` and `description` fields from each `SKILL.md` file, along with the file content, and stores everything in a `Skill` object. `SkillsProvider` loops through these `Skill` objects, formats the `name` and `description` fields like this, and merges them into the agent's system prompt. This keeps the agent aware of available skills without loading their full content upfront. ```python lines.append(" ") lines.append(f" {xml_escape(skill.name)}") lines.append(f" {xml_escape(skill.description)}") lines.append(" ") ``` `SkillsProvider` also adds two methods to the agent through context: `load_skill` and `read_skill_resource`. When the agent decides which skill it needs based on the user's request, it calls `load_skill` to look up the matching `Skill` object by name and loads its full content into the context. If a skill's content references extra resource files like `references/search.md`, the agent can call `read_skill_resource` to load those files. Here's the full workflow: ![A diagram illustrating the workflow of SkillsProvider.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-1.png) A diagram illustrating the workflow of SkillsProvider. Image by Author This design follows the progressive disclosure principle defined by [agentskills.io](https://agentskills.io/home?ref=dataleadsfuture.com). Skill content loads into the agent's context gradually, only when needed. No context explosion, no wasted tokens. --- ## Agent Skills for Enterprise Systems Alright, enough theory. Let's get into today's main topic: **how to use Agent Skills in enterprise-grade agentic systems**. ### Load skills from external systems in real time What if business users write their skills through a cloud-based web page and save them to a database? How do you handle that? We need a new approach to sync and apply Agent Skills in real time. As I covered earlier, when `SkillsProvider` initializes, it loads all `SKILL.md` files from the input paths into an in-memory list of `Skill` objects. Besides the file system approach, `SkillsProvider` also supports Code Defined Skills, where you write skill content directly in code: ```python from pathlib import Path from agent_framework import Skill, SkillsProvider my_skill = Skill( name="my-code-skill", description="A code-defined skill", content="Instructions for the skill.", ) ``` Then pass it to `SkillsProvider` through the `skills` attribute: ```python skills_provider = SkillsProvider( skill_paths=Path(__file__).parent / "skills", skills=[my_skill], ) ``` This opens the door to managing and loading skills from a database. But the original `SkillsProvider` class only accepts skills at initialization time. We want to load skills dynamically while the agent system is running, so we need to extend `SkillsProvider`. After reading the source code, I found that every class extending `BaseContextProvider` has a `before_run` method that gets called when the agent calls `run`. We can load the latest skills from the database before `before_run` executes, then update `SkillsProvider`'s `self._skills` list and refresh the skills description in `instructions`. What I need is a hook method. Every time before `before_run` runs, this hook fetches the latest skills. All I need to do is put the database fetching logic inside this hook. ![The workflow for loading skills from the database in real time.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-2.png) The workflow for loading skills from the database in real time. Image by Author The simplest way to give `SkillsProvider` this hook is to build an `UpdatableSkillsProvider` subclass. This subclass accepts a `skills_updater` parameter at initialization: ```python class UpdatableSkillsProvider(SkillsProvider): def __init__( self, skill_paths: str | Path | Sequence[str | Path] | None = None, *, skills_updater: Callable[[], Awaitable[Sequence[Skill]]] | None = None, **kwargs ): super().__init__( skill_paths=skill_paths, **kwargs, ) self._skills_updater = skills_updater ... ``` `UpdatableSkillsProvider` calls the hook through a private `_update` method, which also updates `self._skills` and the agent's system prompt. Then `before_run` calls `_update` to keep skills fresh in real time: ```python class UpdatableSkillsProvider(SkillsProvider): ... async def _update(self) -> None: if self._skills_updater is None: return try: new_skills = await self._skills_updater() for skill in new_skills: self._skills[skill.name] = skill has_scripts = any(s.scripts for s in self._skills.values()) self._instructions = _create_instructions( prompt_template=self._instruction_template, skills=self._skills, include_script_runner_instructions=has_scripts, ) self._tools = self._create_tools( include_script_runner_tool=has_scripts, require_script_approval=self._require_script_approval, ) except Exception as exc: logger.exception("Failed to update skills: %s", exc) @override async def before_run( self, *, **kwargs ) -> None: await self._update() await super().before_run( **kwargs ) ``` Let's write a `get_latest_skills` hook to simulate loading the latest skills from a database: ```python @lru_cache async def get_latest_skills() -> list[Skill]: """ Pseudocode. In this hook method, you can read the skills text from the database and dynamically build Skill objects. :return: """ code_style_skill = Skill( name="code-style", description="Coding style guidelines and conventions for the team", content=dedent("""\ Use this skill when answering questions about coding style, conventions, or best practices for the team. """), ) return [code_style_skill] ``` Call the agent's `run` method, then check in MLFlow whether the skills loaded by `get_latest_skills` show up in the agent's system prompt: ![The skills loaded from the database have been updated into the agent's system prompt. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-3.png) The skills loaded from the database have been updated into the agent's system prompt. Image by Author The hook method works. We can now load skills from a database in real time. ### Run scripts from skills safely inside containers As of the latest version, Microsoft Agent Framework can't run Python scripts locally or inside containers. But most skills guide the agent through business logic using scripts, so we need to give the agent the ability to run those scripts in a code interpreter. As the predecessor to MAF, Autogen provided a way to run Python scripts inside Docker containers. You can learn about that in this article: [Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent SkillsUse Jupyter code executor to help your agent finish tasks in a smarter way![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-67.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-14.webp)](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) We need something like Autogen's `DockerCommandLineCodeExecutor` for Agent Framework. With the help of AI coding tools, building a code executor for Agent Framework isn't hard. (You can find it in the source code repo at the end of the article.) ```python code_executor = DockerCommandLineCodeExecutor( image="python-code-sandbox", work_dir=work_dir, delete_tmp_files=True, environment={ "TAVILY_API_KEY": os.environ.get("TAVILY_API_KEY"), } ) ``` To keep LLM calls simple, we also need an object-oriented `CodeExecutionTool`: ```python class CodeExecutionTool: """Tool for executing code using a CodeExecutor.""" def __init__(self, executor: CodeExecutor) -> None: self._executor = executor async def execute_code(self, code: str, language: Literal["python", "sh"] = "python") -> str: result = await self._executor.execute_code_blocks( [CodeBlock(code=code, language=language)], CancellationToken(), ) return result.output ``` Next, initialize an `execute_code` tool and wire it up to the agent at initialization: ```python code_tool = CodeExecutionTool(code_executor).execute_code ``` In MLflow, you can see that when the agent needs to search the web, it generates Python code based on the skill's instructions and sends it to the container for execution: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-4.png) The agent generated code based on the skill's instructions and executed it in a container. Image by Author This approach not only lets the agent run code defined in skills, but also keeps that execution safe inside a container. Of course, in a server-side deployment, you'd send code to a centralized Jupyter kernel environment for execution. But that's a whole other story. You can dig into that in my other articles. [How I Crushed Advent of Code And Solved Hard Problems Using Autogen Jupyter Executor and Qwen3A detailed guide on using AI to generate code and solve puzzles automatically and quickly![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-68.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/jupyter_executor_cover-2.webp)](https://www.dataleadsfuture.com/how-i-crushed-advent-of-code-and-solved-hard-problems-using-autogen-jupyter-executor-and-qwen3/) ### Reduce context length even further Agent Skills uses progressive disclosure to keep irrelevant skill content from eating up your context window. But as the conversation or task moves forward, skill content that was loaded into earlier messages will still pile up in the context over time. Agent systems today have several context pruning techniques available. Context trimming and context compression, both common in desktop agents, work really well. ![The difference between context pruning and context compression.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-5.png) The difference between context pruning and context compression. Image by Author Beyond those two, today I want to share a context engineering technique I discovered at work that fits Agent Skills even better. As you know, in enterprise scenarios, loading a skill usually means running one atomic workflow: researching a topic through web search? Sure. Running a SWOT analysis on a company and writing a report? No problem. These workflows all share one thing in common. You give the agent the right input, then wait for it to return an output. Which skill the agent loaded, and how it worked through the task — I honestly don't care. I wouldn't even mind if the agent unloaded the skill after finishing to save tokens. That sounds a lot like how a function works. So, can we use an agent with skills loaded as a tool for another agent? Absolutely. That's exactly what I do. Microsoft Agent Framework has a method on Agent called `as_tool`. It turns an agent into a function-callable tool. So I designed a main agent. The main agent takes user requests and generates the right response to return. The agent with Agent Skills loading capability turns itself into a tool for the main agent using `as_tool`. ```python agent = chat_client.as_agent( name="Assistant", instructions=dedent(""" You're a smart little helper who, for each user request, picks the right task description to call a tool, gets the answer, and then delivers the final result. """), tools=[skills_agent.as_tool()], ) ``` The skills agent's workflow stays the same. It loads the right skill based on the task description, generates and runs code, then returns the result. ![The skill agent is provided as a tool for the main agent to call.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-6.png) The skill agent is provided as a tool for the main agent to call. Image by Author But the main agent is different. Its context only holds user messages, the message calling the skills agent tool, and the final response. No skill-related content at all. The main agent's context stays clean, and even after running for a long time, it won't interfere with the LLM. ![Keep the main agent's context clean by loading skills into the sub-agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/03/image-7.png) Keep the main agent's context clean by loading skills into the sub-agent. Image by Author There's a nice bonus too. LLMs know what they want better than humans do, so before the main agent calls the skills agent, it rewrites the user's task into something more precise. This helps the skills agent execute more accurately. --- ## Conclusion That's everything I have for you today on Agent Skills for enterprise agent systems. Unlike desktop agents, enterprise agent systems run on cloud servers. There's no way to update an agent's skills through the file system in real time without downtime. So I went with a targeted approach. This approach lets users write skill content through a web interface and save it to a database, while agents read the latest skills in real time and sync them across server nodes. I used the latest version of Microsoft Agent Framework to build this, but you can use any other framework. The principles are the same. I also covered how to run scripts the agent generates from skills inside containers, which is much safer than running scripts directly on a desktop system. I shared a context management approach I found at work that works especially well for skills-based agents. The Microsoft Agent Framework API is still a bit unstable. If anything is unclear, feel free to leave a comment, and I'll get back to you as soon as I can. Thanks for subscribing! Share this with your friends if you think it might help someone else. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Source Code Get the source code for this article. ****Start using it right away and save yourself a ton of trial and error time.** [Grab the Source Code ](#/portal/signup) ### Advanced RedisVL Long-term Memory Tutorial: Using an LLM to Extract Memories URL: https://www.dataleadsfuture.com/advanced-redisvl-long-term-memory-tutorial-using-an-llm-to-extract-memories/ Last updated: 2026-07-24T12:20:25.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* ## Introduction In this weekend note, we keep talking about how to build long-term memory for an agent with RedisVL. When we build a long-term memory module for an agent, we need to care about two points most: - After long running, will the saved memories grow too large and cause context explosion? - How do we recall the memories that matter most to the current context? We will solve these two problems today. **TLDR**: In this hands-on tutorial, we first use an LLM to extract information from user messages that has value for later chats. Then we store that as long-term memory in RedisVL. When needed, we search related memories with semantic search. With this setup, the agent understands the past context of the user and gives more accurate answers. ![After using an LLM to extract memories, how my long‑term memory module runs.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image.png) After using an LLM to extract memories, how my long‑term memory module runs. Image by Author With this kind of long-term memory, we do not worry about memory explosion after long running. We also do not worry that unrelated memories will hurt LLM responses. You can get all the source code at the end of this post. ### Why do we do this? In the last hands-on post, I shared how to build short-term and long-term memory for an agent with RedisVL: [Build Long-Term and Short-Term Memory for Agents Using RedisVLPros and cons analysis based on real-world practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-62.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-11.webp)](https://www.dataleadsfuture.com/build-long-term-and-short-term-memory-for-agents-using-redisvl/) The short-term part works very well. RedisVL API feels much simpler than the raw Redis API. I can write that code with little effort. The long-term part does not work. We follow the official example and store user queries and LLM responses in RedisVL. When chat continues, RedisVL keeps pulling repeated queries or unrelated answers via semantic search. This troubles the LLM a lot and blocks the chat from going on. ### Can we avoid RedisVL? Your boss will not agree. There is already Redis in your stack. Why do you still want to install mem0 or other open source tools? How about extra cost? This is real life. So we still need to make RedisVL work. But I do not want to run around like a headless fly. Before that I want to see how humans handle memory. --- ## How Humans Handle Memory ### What deserves memory First, we need to know one thing. Only information that centers on me and links to me tightly deserves a sticky note. So what information about me do I want to write down? - Preference settings such as tools I like, languages I use, my schedule, and my tone when I talk - Stable personal info such as my role, my time zone, and my daily habits - Goals and decisions, such as chosen options, plans, and sentences that start with “I decide to...” - Key milestones such as job change, moving, deadlines, and product launches - Work and project context, such as project names, stakeholders, needs, and status like “done/next” step. - Repeated pain points or strong views that will change LLM advice later - Things I say with “remember this ...” or “do not forget ...” ### What does not deserve memory I do not plan to store any LLM answer. LLM answers to the same question will change with context. So LLM answers in long-term memory do not help much. Besides LLM answers, I also do not want to keep these: - One-time small things that likely will not matter later - Very sensitive personal data such as health diagnosis, exact address, government IDs, passwords, bank accounts - Things I clearly ask not to remember - Things I already wrote down on the sticky note --- ## Design a Prompt for LLM Memory Extraction Now we know how humans handle memory. Next, I want to build an agent that follows the same rules and extracts memories from my daily chats. The key lives in the `system prompt`. I need to describe all rules in the system prompt. Then I ask the agent to follow these rules with very high consistency. In the past, I might have tried some “write 1000-line prompt” challenge. Now I do not need that. I just open any LLM client, paste these rules, then ask the LLM to help me write a `system prompt`. This takes less than one minute. After a few tries, I pick one I like. Here is that `system prompt`: ```latex Your job: Based only on the user’s current input and the existing related memories, decide whether you need to add a new “long-term memory,” and if needed, **extract just one fact**. You do not talk to the user. You only handle memory extraction and deduplication. --- ### 1. Core principles 1. Only save information that **will likely be useful in the future**. 2. **At most one fact per turn**, and it must clearly appear in the current input. 3. **Never invent or infer anything**. You can only restate or lightly rephrase what the user has explicitly said. 4. If the current input has nothing worth keeping, or the information is already in the related memories, then do not add a new memory. --- ### 2. What counts as “long-term memory” Only consider the categories below, and decide whether the information has long-term value: ... ``` Due to space, I only show part of the prompt here. You can get the full prompt from the source code at the end. --- ## Build a ContextProvider for Long-term Memory After we finish the memory extraction rules, we start to build the long-term memory module for the agent. For future use, I still pick Microsoft Agent Framework MAF. It gives a `ContextProvider` feature that lets us plug long-term memory into the agent in a simple way. Of course the principle of long-term memory stays the same. You can use any agent framework you like and build your own memory module. Or you can ignore frameworks and first build storage and retrieval of memories, and then call them through function calls. That is fine. ### Run memory extraction in sequence In the last post, I already built a long-term memory module with `ContextProvider`. The new version looks similar. But this time, we use an LLM to extract memories. So after we set up `ContextProvider`, we first use the `system prompt` to build a memory extraction agent. If you do not know how to use `ContextProvider` yet, I suggest you read my last post again. That post explains `ChatMessageStore` and `ContextProvider` in Microsoft Agent Framework in detail: [Build Long-Term and Short-Term Memory for Agents Using RedisVLPros and cons analysis based on real-world practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-63.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-12.webp)](https://www.dataleadsfuture.com/build-long-term-and-short-term-memory-for-agents-using-redisvl/#/portal) To avoid too much unrelated data from Redis, I set the `distance_threshold` value pretty small. But not too small. If too small, then it loses meaning. You can pick the value you like. ```python class LongTermMemory(ContextProvider): def __init__( self, thread_id: str | None = None, session_tag: str | None = None, distance_threshold: float = 0.3, context_prompt: str = ContextProvider.DEFAULT_CONTEXT_PROMPT, redis_url: str = "redis://localhost:6379", embedding_model: str = "BAAI/bge-m3", llm_model: str = Qwen3.NEXT, llm_api_key: str | None = None, llm_base_url: str | None = None, ): ... self._init_extractor() def _init_extractor(self): with open("prompt.md", "r", encoding="utf-8") as f: system_prompt = f.read() self._extractor = OpenAILikeChatClient( model_id=self._llm_model, ).as_agent( name="extractor", instructions=system_prompt, default_options={ "response_format": ExtractResult, "extra_body": {"enable_thinking": False} }, ) ``` Next we implement the `invoking` method. This method runs before the user agent calls the LLM. In this method, we extract and store long-term memory To make the logic clear, I first implement the `invoking` method in order, as in this diagram: ![Run the memory extraction logic step by step in order.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-1.png) Run the memory extraction logic step by step in order. Image by Author When a new user request comes into `ContextProvider`, we first search RedisVL with semantic search for the most similar memories. ```python class LongTermMemory(ContextProvider): ... async def invoking( self, messages: ChatMessage | MutableSequence[ChatMessage], **kwargs: Any ) -> Context: if isinstance(messages, ChatMessage): messages = [messages] prompt = "\n".join([m.text for m in messages]) line_sep_memories = self._get_line_sep_memories(prompt) ... def _get_line_sep_memories(self, prompt: str) -> str: context = self._semantic_store.get_relevant(prompt, role="user", session_tag=self._session_tag) line_sep_memories = "\n".join([f"* {str(m.get("content", ""))}" for m in context]) return line_sep_memories ``` Next, we send these existing memories plus the user request to the memory extraction agent. That agent first checks if anything is worth saving according to the rules. Then it extracts a new helpful memory from the user request and saves it into RedisVL. ```python class LongTermMemory(ContextProvider): ... async def invoking( self, messages: ChatMessage | MutableSequence[ChatMessage], **kwargs: Any ) -> Context: ... await self._save_memory(messages, line_sep_memories) ... async def _save_memory( self, messages: ChatMessage | MutableSequence[ChatMessage], relevant_memory: str | None = None, ) -> None: detect_messages = ( [ ChatMessage(role=Role.USER, text=f"Existing related memories:\n\n{relevant_memory}"), ] + list(messages) if relevant_memory.strip() else list(messages) ) response = await self._extractor.run(detect_messages) extract_result: ExtractResult = cast(ExtractResult, response.value) if extract_result.should_write_memory: self._semantic_store.add_messages( messages=[ {"role": "user", "content": extract_result.memory_to_write} ], session_tag=self._session_tag, ) ``` Last, we put the memories from RedisVL into `Context` as extra context. These memory messages get merged into the history of the real chat agent. They give the chat agent extra background to produce answers. ```python class LongTermMemory(ContextProvider): ... async def invoking( self, messages: ChatMessage | MutableSequence[ChatMessage], **kwargs: Any ) -> Context: ... return Context(messages=[ ChatMessage(role="user", text=f"{self._context_prompt}\n{line_sep_memories}") ] if len(line_sep_memories)>0 else None) ``` Now we build a simple chat agent to test the new long-term memory module ```python agent = OpenAILikeChatClient( model_id=Qwen3.MAX ).as_agent( name="assistant", instructions="You are a helpful assistant.", context_provider=LongTermMemory(), ) async def main(): thread = agent.get_new_thread() while True: user_input = input("\nUser: ") if user_input.startswith("exit"): break stream = agent.run_stream(user_input, thread=thread) print("Assistant: \n") async for event in stream: print(event.text, end="", flush=True) print("\n") if __name__ == "__main__": asyncio.run(main()) ``` Now we chat with the agent and see how it works: ![The yellow parts are memories fetched from RedisVL, and the green parts are memories extracted by the LLM.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-2.png) The yellow parts are memories fetched from RedisVL, and the green parts are memories extracted by the LLM. Image by Author From MLFlow we see that the retrieved memories go in as a separate message in the chat history: ![The retrieved memory will be added to the chat history as a separate message.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-3.png) The retrieved memory will be added to the chat history as a separate message. Image by Author Then we check Redis and see what memories we saved: ![The info that the LLM pulls out from the user’s past conversations will be saved into Redis as valuable memories.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-4.png) The info that the LLM pulls out from the user’s past conversations will be saved into Redis as valuable memories. Image by Author We can see that as the chat goes on, the new long-term memory module no longer stores and retrieves all chat history without filter. It keeps and retrieves only memories that matter to the user, and these memories give strong help in later chats. ### Use concurrency to speed up Everything looks fine except for the part where we use an LLM to extract memories that deserve saving. The largest delay in an agent often comes from LLM calls. Now we add one more LLM call. We also need to wait for the LLM to decide whether to save memory before we go on with the real chat. We can add some logs and see how much delay we add: ![Every time we use the LLM to pull out memories, it takes a bit more than one second.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-5.png) Every time we use the LLM to pull out memories, it takes a bit more than one second. Image by Author We add more than one second per chat turn. One way to optimize is to use a smaller model like `qwen3-8b`. But the gain stays small. We save little time and hurt memory quality due to the smaller model. Today I use a different way. I use concurrent programming so that the LLM call for memory extraction and the LLM call for user reply run at the same time. Let us see the result after that change: ![The memory retrieval step basically doesn’t take any time at all, so how did I pull that off?](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-6.png) The memory retrieval step basically doesn’t take any time at all, so how did I pull that off? Image by Author The time cost for extraction and storage of memory becomes almost nothing while the effect stays the same. How do we reach that? If you built multi-agent workflows with LangGraph or LlamaIndex, you have likely seen the fan-out idea. It lets many nodes run at the same time, and then you take the final result. The base idea uses the `asyncio` module in Python. You often see `async` and `await` when you write agent code. I wrote many posts in the past about `asyncio` and concurrency: [Use These Methods to Make Your Python Concurrent Tasks Perform BetterBest practices for asyncio.gather, asyncio.as\_completed, and asyncio.wait![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-64.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/0_Eel1JPfhnioG71TL.jpg)](https://www.dataleadsfuture.com/use-these-methods-to-make-your-python-concurrent-tasks-perform-better/) In short, when you face delays because of long IO calls, you can use concurrent programming. **Note:** If you use `mlflow.openai.autolog()` to trace LLM calls, you may see that concurrent runs stop working. I still do not know why. I suggest you comment out MLFlow parts before you go on. ![By running the memory retrieval process with concurrent programming.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/02/image-7.png) By running the memory retrieval process with concurrent programming. Image by Author Back to our current case. We see that `_save_memory` is an `async` method. That means we can run it with concurrency. How do we do that? Very simple. When we call `_save_memory`, we use `asyncio.create_task` to build a new concurrent task. That is all. ```python class LongTermMemory(ContextProvider): ... async def invoking( self, messages: ChatMessage | MutableSequence[ChatMessage], **kwargs: Any ) -> Context: ... asyncio.create_task(self._save_memory(messages, line_sep_memories)) ... ``` Since real user chat often takes more time than memory extraction, we do not need to wait for that task in code. We only need to create the task. With that, we add a memory extraction module that does not bring much extra delay to the agent system. --- ## Conslusion Redis now serves as standard infra for many companies. With RedisVL, it can cache and search information by semantics. This makes it easier to build short-term and long-term memory on top of Redis. But if you build long-term memory with RedisVL API directly, you may not see good results. The system has no “brain” to judge which information deserves long-term storage and keeps value over time. So in this tutorial, I first use an agent to extract useful memories and then write them into RedisVL. This improves the value of saved information. Long-term memory works much better now and fills the gap in my last post. I also share a short guide on how to use concurrent programming so that many LLM calls run at the same time. This cuts system delay by a lot. If you like concurrency, you can read my old posts. Thanks for reading. If you have any questions or ideas, leave me a note. I will reply as soon as I can. Do not forget to subscribe to my blog and follow my new work in AI applications. Also share this post with your friends. It may help more people. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM AI Developer Professional Certificate**](https://imp.i384100.net/B5eg04?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- Get the source code for this article. ****Start using it right away and save yourself a ton of trial and error time.** [Grab the Source Code ](#/portal/signup) ### Build Long-Term and Short-Term Memory for Agents Using RedisVL URL: https://www.dataleadsfuture.com/build-long-term-and-short-term-memory-for-agents-using-redisvl/ Last updated: 2026-04-22T08:18:47.000Z ## Introduction For this weekend note, I want to share some tries I made using RedisVL to add short-term and long-term memory to my agent system. TLDR: RedisVL works pretty well for short-term memory. It feels a bit simpler than using the traditional Redis API. For long-term memory with semantic search, the experience is not good. I do not recommend it. ### Why RedisVL? Big companies like to use mature infrastructure to build new features. We know mem0 and Graphiti are good open source software for long-term agent memory. But companies want to stay safe. Building new infrastructure costs money. It is unstable. It needs people who know how to run it. So when Redis launched RedisVL with vector search, we naturally wanted to try it first. You can connect it to existing Redis clusters and start using it. That sounds nice. But is it really nice? We need to try it for real. Today I will cover how to use `MessageHistory` and `SemanticMessageHistory` from RedisVL to add short-term and long-term memory to agents built on the Microsoft Agent Framework. You can find the source code at the end of this article. --- ## Preparation ### Install Redis If you want to try it locally, you can install a Redis instance with Docker. ```shell docker run -d --name redis -p 6379:6379 -p 8001:8001 redis/redis-stack:latest ``` Cannot use Docker Desktop? See my other article. [A Quick Guide to Containerizing Agent Applications with PodmanAlternative solutions compatible with Docker SDK![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-60.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/podman.webp)](https://www.dataleadsfuture.com/a-quick-guide-to-containerizing-agent-applications-with-podman/) The Redis instance will listen on ports 6379 and 8001\. Your RedisVL client should connect to `redis://localhost:6379`. You can visit `http://localhost:8001` in the browser to open the Redis console. ### Install RedisVL Install RedisVL with pip. ```shell pip install redisvl ``` After installation, you can use the RedisVL CLI to manage your indexes and keep your testing neat. ```shell rvl index listall ``` --- ## Implement Short-Term Memory Using MessageHistory There are lots of “How to” RedisVL articles online, so let’s start straight from Microsoft Agent Framework and see how to use `MessageHistory` for short-term memory. As in the official tutorial, you should implement a `RedisVLMessageStore` based on `ChatMessageStoreProtocol`. ```python class RedisVLMessageStore(ChatMessageStoreProtocol): def __init__( self, thread_id: str = "common_thread", top_k: int = 6, session_tag: str | None = None, redis_url: str | None = "redis://localhost:6379", ): self._thread_id = thread_id self._top_k = top_k self._session_tag = session_tag or f"session_{uuid4()}" self._redis_url = redis_url self._init_message_history() ``` In `__init__` you should note two parameters. - `thread_id` is used for the `name` parameter when creating `MessageHistory`. I like to bind it to the agent. Each agent gets a unique `thread_id`. - `session_tag` lets you set a tag for each user so different sessions do not mix. The protocol asks us to implement two methods `list_messages` and `add_messages`. - `list_messages` runs before the agent calls the LLM. It gets all available chat messages from the message store. It takes no parameters, so it cannot support long-term memory. More on that later. - `add_messages` runs after the agent gets the LLM’s reply. It stores new messages into the message store. Here is how the message store works. ![The calling order of message store in the agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-5.png) The calling order of message store in the agent. Image by Author So in `list_messages` and `add_messages` we just use RedisVL’s `MessageHistory` to do the job. `list_messages` below uses `get_recent` to get `top_k` recent messages and turns them into `ChatMessage`. ```python class RedisVLMessageStore(ChatMessageStoreProtocol): ... async def list_messages(self) -> list[ChatMessage]: messages: list[dict[str, str]] = self._message_history.get_recent( top_k=self._top_k, session_tag=self._session_tag, ) return [self._back_to_chat_message(message) for message in messages] ``` `add_messages` turns the `ChatMessage` into Redis messages and calls `add_messages` to store them. ```python class RedisVLMessageStore(ChatMessageStoreProtocol): ... async def add_messages(self, messages: Sequence[ChatMessage]): messages = [self._to_redis_message(message) for message in messages] self._message_history.add_messages( messages, session_tag=self._session_tag ) ``` That is short-term memory done with RedisVL. You may also implement `deserialize`, `serialize` and `update_from_state` for saving and loading the memory, but it is not important now. See the full code at the end. ### Test RedisVLMessageStore Let’s build an agent and test the message store. ```python agent = OpenAILikeChatClient( model_id=Qwen3.NEXT ).create_agent( name="assistant", instructions="You're a little helper who answers my questions in one sentence.", chat_message_store_factory=lambda: RedisVLMessageStore( session_tag="user_abc" ) ) ``` Now a console loop for multi-turn dialog. Remember, Microsoft Agent Framework does not support short-term memory unless you use an `AgentThread` and pass it to `run`. ```python async def main(): thread = agent.get_new_thread() while True: user_input = input("User: ") if user_input.startswith("exit"): break response = await agent.run(user_input, thread=thread) print(f"\nAssistant: {response.text}") thread.message_store.clear() ``` `AgentThread` when created calls the factory method to build the `RedisVLMessageStore`. To check if the store works, we can use `mlflow.openai.autolog()` to see if messages sent to the LLM contain historical messages. ```python import mlflow mlflow.set_tracking_uri(os.environ.get("MLFLOW_TRACKING_URI")) mlflow.set_experiment("Default") mlflow.openai.autolog() ``` ![You can see that the conversation comes with a complete history of messages.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-6.png) You can see that the conversation comes with a complete history of messages. Image by Author See my other article for using MLFlow to track LLM calls. [Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracing TutorialEnhance your multi-agent application’s observability, explainability and Traceability![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-61.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Tracing-with-MLflow-3.x-1-1.webp)](https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/) Let’s open the Redis console to see the cache. ![How the cache is stored in Redis.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-7.png) How the cache is stored in Redis. Image by Author As you can see, after using `MessageHistory` as MAF's message store, we can implement multi-turn conversations with historical messages. With `thread_id` and `session_tag` parameters, we can also implement the feature that lets users switch between multiple conversation sessions, like in popular LLM chat applications. Feels simpler than the official `RedisMessageStore` solution right? --- ## Implement Long-Term Memory Using SemanticMessageHistory `SemanticMessageHistory` is a subclass of `MessageHistory`. It adds a `get_relevant` method for vector search. Example: ```python prompt = "what have I learned about the size of England?" semantic_history.set_distance_threshold(0.35) context = semantic_history.get_relevant(prompt) for message in context: print(message) ``` ```latex Batches: 100%|██████████| 1/1 [00:00<00:00, 56.30it/s] {'role': 'user', 'content': 'what is the size of England compared to Portugal?'} ``` Compared to `MessageHistory` the big thing here is that we can get the most relevant historical messages based on the user request. You might think that if `MessageStore` short-term memory is nice, then `SemanticMessageHistory` with semantic search must be even better. From my experience, this is not the case. From my test results, it is not like that. Let’s now make a long-term memory adapter for Microsoft Agent Framework using `SemanticMessageHistory` and see the result. In my latest research, I’m going to bring in an LLM to handle memory retrieval so long‑term memory can become more interconnected: [Advanced RedisVL Long-term Memory Tutorial: Using an LLM to Extract MemoriesBuilding an intelligent, context-aware memory system![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-65.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-13.webp)](https://www.dataleadsfuture.com/advanced-redisvl-long-term-memory-tutorial-using-an-llm-to-extract-memories/) ### Use SemanticMessageHistory in Microsoft Agent Framework Earlier I said `list_messages` in `ChatMessageStoreProtocol` has no parameters, so we cannot search history. Thus, we cannot use `MessageStore` for long-term memory. Microsoft Agent Framework has a `ContextProvider` class. From its name, it is for context engineering. So we should build long-term memory on this class. ```python class RedisVLSemanticMemory(ContextProvider): def __init__( self, thread_id: str | None = None, session_tag: str | None = None, distance_threshold: float = 0.3, redis_url: str = "redis://localhost:6379", embedding_model: str = "BAAI/bge-m3", embedding_api_key: str | None = None, embedding_endpoint: str | None = None, ): self._thread_id = thread_id or "semantic_thread" self._session_tag = session_tag or f"session_{uuid4()}" self._distance_threshold = distance_threshold self._redis_url = redis_url self._embedding_model = embedding_model self._embedding_api_key = embedding_api_key or os.getenv("EMBEDDING_API_KEY") self._embedding_endpoint = embedding_endpoint or os.getenv("EMBEDDING_ENDPOINT") self._init_semantic_store() ``` `ContextProvider` has two methods `invoked` and `invoking`. - `invoked` runs after LLM call. It stores the latest messages in RedisVL. It has both `request_message` and `response_messages` parameters but stores them separately. - `invoking` runs before LLM call. It uses the user’s current input to search for relevant history in RedisVL and returns a `Context` object. The `Context` object has three variables. - `instructions` string. The agent adds this to the system prompt. - `messages` list. Put history messages found in long-term memory here. - `tools` list for functions. The agent adds these tools to its `ChatOptions`. ![The purpose of the three types of messages retrieved.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-8.png) The purpose of the three types of messages retrieved. Image by Author Since we want to use vector search to get relevant history, we put those messages in `messages`. The order between `MessageStore` messages and `ContextProvider` messages matters. Here is the order of their calls. ![The calling order of long-term and short-term memory in the agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-9.png) The calling order of long-term and short-term memory in the agent. Image by Author ### Setting up a TextVectorizer Semantic vector search needs embeddings. We must set up a vectorizer. In `__init__` besides `thread_id` and `session_tag` we set the embedding model info. ```python class RedisVLSemanticMemory(ContextProvider): ... def _init_semantic_store(self) -> None: if not self._embedding_api_key: vectorizer = HFTextVectorizer( model=self._embedding_model, ) else: vectorizer = OpenAITextVectorizer( model=self._embedding_model, api_config={ "api_key": self._embedding_api_key, "base_url": self._embedding_endpoint } ) self._semantic_store = SemanticMessageHistory( name=self._thread_id, session_tag=self._session_tag, distance_threshold=self._distance_threshold, redis_url=self._redis_url, vectorizer=vectorizer, ) ``` I can choose a server-hosted embedding model with OpenAI API or a local HuggingFace model, depending on whether `embedding_api_key` is set. ### Implement invoked and invoking methods `invoked` is easy. As said `SemanticMessageHistory` stores request and response separately. I merge them into one list, then call `add_messages`. ```python class RedisVLSemanticMemory(ContextProvider): ... async def invoked( self, request_messages: ChatMessage | Sequence[ChatMessage], response_messages: ChatMessage | Sequence[ChatMessage] | None = None, invoke_exception: Exception | None = None, **kwargs: Any, ) -> None: if isinstance(request_messages, ChatMessage): request_messages = [request_messages] if isinstance(response_messages, ChatMessage): response_messages = [response_messages] chat_messages = request_messages + response_messages messages = [self._to_redis_message(message) for message in chat_messages] self._semantic_store.add_messages( messages=messages, session_tag=self._session_tag, ) ``` `invoking` below: ```python class RedisVLSemanticMemory(ContextProvider): ... async def invoking( self, messages: ChatMessage | MutableSequence[ChatMessage], **kwargs: Any ) -> Context: if isinstance(messages, ChatMessage): # 1 messages = [messages] prompt = "\n".join([message.text for message in messages]) context = self._semantic_store.get_relevant( prompt=prompt, raw=True, session_tag=self._session_tag, ) context = sorted(context, key=lambda m: m['timestamp']) # 2 relevant_messages = [self._back_to_chat_message(message) for message in context] print([m.text for m in relevant_messages]) return Context(messages=relevant_messages) # 3 ``` Points to note. - The `messages` parameter may be a list for multi-modal input. Merge all text. - Since messages are stored separately, I need to sort them by timestamp to keep order. - Put the retrieved messages into `Context.messages` so they go to the end of the current chat messages. ### Test semantic memory Unlike message store, we can set `ContextProvider` directly in the agent. ```python memory_provider = RedisVLSemanticMemory( session_tag="user_abc", distance_threshold=0.3, ) agent = OpenAILikeChatClient( model_id=Qwen3.NEXT ).create_agent( name="assistant", instructions="You're a little helper who answers my questions in one sentence.", context_providers=memory_provider, ) ``` Now a `main` with a `thread` instance to keep short-term memory while testing multi turn dialog. ```python async def main(): thread = agent.get_new_thread() while True: user_input = input("User: ") if user_input.startswith("exit"): break response = await agent.run(user_input, thread=thread) print(response.text) memory_provider.clear() ``` Test result: ![The distance_threshold is too high, causing irrelevant messages to be retrieved.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-10.png) The `distance_threshold` is too high, causing irrelevant messages to be retrieved. Image by Author It seems the default value of `distance_threshold` 0.3 is too high. Let's set it lower: ```python memory_provider = RedisVLSemanticMemory( session_tag="user_abc", distance_threshold=0.2, ) ``` Test again: ![Only request messages were retrieved, not response messages.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-11.png) Only request messages were retrieved, not response messages. Image by Author Lower threshold stops unrelated messages. But since requests and responses are stored separately, only requests are found. ContextProvider puts retrieved messages at the end of the message list. The LLM may think the user asked two questions. MLFlow shows it. ![Two similar questions were both added to the message list, but without attaching the already provided answers.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-12.png) Two similar questions were both added to the message list, but without attaching the already provided answers. Image by Author This is bad. We care more about the LLM’s answers than the requests. But vector search often finds the questions, not the answers. This just adds useless questions and does not help the LLM answer. Hard to say if the fault is Microsoft Agent Framework or RedisVL. When `ContextProvider`long-term finds related chat messages, they go after the ones from message store. If long-term and short-term messages repeat, they can confuse the LLM. Also, RedisVL not storing requests and responses together is a choice I do not like. LLM responses cost more. In production, a response may involve web search, RAG retrieval, or running code. But vector search finds just the request, not the answer. That is a waste. --- ## Conclusion Today, we tried using RedisVL for short-term and long-term memory in Microsoft Agent Framework and checked the results. RedisVL is very handy for short-term agent memory. It is simpler than using the Redis API. But `SemanticMessageHistory` for semantic search of the user history did not perform well. I explained why. Thanks to the solid Redis infrastructure, semantic caches with RedisVL are simpler than other vector solutions. Next time, I will show you a semantic cache with RedisVL to save big costs for your company. Share your thoughts in the comments. Subscribe to my blog to follow my latest agent app work. And share this article with friends. Maybe it will help more people.😁 --- Here's the source code of this article: [agentic-ai-playground/12\_RedisVL\_Long\_Short\_Memory at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-30.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-28)](https://github.com/qtalen/agentic-ai-playground/tree/main/12%5FRedisVL%5FLong%5FShort%5FMemory?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Microsoft Agent Framework (MAF) Middleware Basics: Add Compliance Fences to Your Agent URL: https://www.dataleadsfuture.com/microsoft-agent-framework-maf-middleware-basics-add-compliance-fences-to-your-agent/ Last updated: 2026-04-22T08:19:22.000Z ## Introduction The recent incident where Tencent’s Yuanbao AI insulted users once again shows that no matter how powerful an agent is, you must set up compliance guardrails and carefully review its inputs and outputs before going live. When we build enterprise-grade agent applications, we often work on a large-scale system involving multiple teams. Beyond the parts handled by our own team, we also need to dynamically plug in features from other teams—such as permissions, logging, billing, and compliance review. These features should not interfere with the agent system’s core capabilities, yet they must be easy to install and remove. In traditional web frameworks like FastAPI, middleware provides exactly this kind of dynamic code injection capability. So, does something similar exist in agent systems? Actually, yes. In today’s tutorial, I’ll show you how I used Microsoft Agent Framework’s middleware and AG-UI features to add compliance review for user inputs into my agent. ![Inducing an agent will be blocked by compliance rules specific to certain business scenarios.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image.png) Inducing an agent will be blocked by compliance rules specific to certain business scenarios. Image by Author This will teach you how to use middleware in enterprise agent applications and give you a first look at using AG-UI for microservice distributed agent development. Let’s start. You can get all the source code at the end. 👇 ***Don’t forget to follow my blog to stay updated on my latest progress in AI application practices.** [Subscribe Now ](#/portal/signup) --- ## System Setup ### Install the latest Microsoft Agent Framework MAF is still updating quickly. Since APIs change a lot, this guide uses the newest version. It’s better to install the prerelease version. ```shell pip install agent-framework --pre ``` Or add the dependency in your `pyproject.toml` file: ```latex "agent-framework-ag-ui>=1.0.0b260127" ``` ### Install Microsoft Agent Framework AG-UI MAF works with AG-UI to support distributed agent development. You’ll need this capability today, so install the latest version of ag-ui; otherwise, APIs won’t match up. ```latex "agent-framework-ag-ui>=1.0.0b260127" ``` After installing the needed Python packages, we can move on. First, let's get a quick background on what middleware is and what it can do. --- ## Quick Intro to MAF Middleware ### What is middleware According to the MAF documentation: > *Middleware in the Agent Framework intercepts, changes, and enhances agent behavior at different execution points. You can use it for logging, security checks, error handling, and result transformation without changing the agent’s or function’s core logic.* That’s what we’ll learn today. ### How middleware works As I said before, MAF middleware uses the chain-of-responsibility pattern. Each piece of logic lives in its own node. Every node knows the next one. When a node finishes running, it passes control to the next node. Here’s a simple example: ```python async def logging_agent_middleware( context: AgentRunContext, next: Callable[[AgentRunContext], Awaitable[None]] ) -> None: print("[Agent] Starting execution") await next(context) print("[Agent] Execution completed") ``` The `next` parameter points to the next node. You can run code before or after calling it. The actual agent logic acts as the last node. After all middleware nodes finish, the agent runs. In MAF, middleware can run in three stages: - Before or after `run` or `run_stream`. - Before or after a function call. - Before or after calling the LLM. ![The Microsoft Agent Framework middleware works at different stages of agent execution. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-1.png) The Microsoft Agent Framework middleware works at different stages of agent execution. Image by Author Now let’s look at how different middleware types work. ### Function-Based middleware If your middleware is simple, like just logging agent runs, use a function-based middleware. You only need a function with two parameters: `context` and `next`. The `context` keeps your runtime info, and `next` calls the next node. MAF uses the type annotation of `context` to tell which stage this code belongs to. For example, if it runs at the agent stage, the type should be `AgentRunContext`: ```python async def logging_agent_middleware( context: AgentRunContext, next: Callable[[AgentRunContext], Awaitable[None]] ) -> None: print("[Agent] Starting execution") await next(context) print("[Agent] Execution completed") ``` For a function call stage, use `FunctionInvocationContext`: ```python async def logging_function_middleware( context: FunctionInvocationContext, next: Callable[[FunctionInvocationContext], Awaitable[None]], ) -> None: print(f"[Function] Calling {context.function.name}") await next(context) print(f"[Function] {context.function.name} completed") ``` And for the chat stage, use `ChatContext`: ```pyhon async def logging_chat_middleware( context: ChatContext, next: Callable[[ChatContext], Awaitable[None]], ) -> None: print(f"[Chat] Sending {len(context.messages)} messages to AI.") await next(context) print(f"[Chat] AI response received.") ``` If you dislike type annotations, you can use decorators. `@agent_middleware` runs at the agent stage: ```python @agent_middleware async def logging_agent_middleware(context, next) -> None: print("[Agent] Starting execution") await next(context) print("[Agent] Execution completed") ``` Then you don’t need to add type annotations anymore. There are also `@function_middleware` and `@chat_middleware` for function calls and chat calls. If your middleware needs to save state or handle more complex logic, function-based won’t be enough. Use class-based middleware. ### Class-Based middleware Class-based middleware organizes code with object-oriented methods. That lets middleware remember state and handle tricky logic. A class-based middleware must meet two rules: 1. Inherit from the right base class: `AgentMiddleware`, `FunctionMiddleware`, or `ChatMiddleware`. 2. Have a `process` method with the same parameters as the function-based ones. They use the same contexts. Here’s an example for a middleware class that runs at the function call stage: ```python class LoggingFunctionMiddleware(FunctionMiddleware): async def process( self, context: FunctionInvocationContext, next: Callable[[FunctionInvocationContext], Awaitable[None]] ) -> None: print(f"[Function Class] Calling {context.function.name}") await next(context) print(f"[Function Class] {context.function.name} completed.") ``` Just make sure to pair the right base class with the right context type. The others follow the same rule. ### How to use middleware There are three stages for middleware and three ways to build it. Let’s put that in one grid chart to see how they connect. ![Use a grid chart to describe the implementations of different middleware.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-2.png) Use a grid chart to describe the implementations of different middleware. Image by Author The framework now only supports passing middleware when creating the agent: ```python agent = chat_client.as_agent( name="assistant", instructions="You are a helpful assistant", tools=[get_weather], middleware=[ logging_agent_middleware, LoggingFunctionMiddleware(), logging_chat_middleware, blocking_middleware, logging_function_middleware, ] ) ``` You can mix all nine types freely. But note that **only the last function middleware you add actually works right now. I’m not sure if that’s a bug, but we’ll find out later.** --- ## Project Practice: Add Compliance Check to Your Agent Now let’s get hands-on. I’ll show how to use MAF middleware to add compliance checking to an agent. ### Why add compliance checks Every LLM already has basic compliance setups built in based on local laws. When companies self-host LLMs, they also add custom checks in frameworks like vLLM. But those only watch the model’s input or output. Now that agents are everywhere, we also need checks at the agent level: preventing prompt injection, checking MCP permissions, and so on. Middleware makes this possible. In today’s demo, we’ll review every user message to make sure no one tries to make our finance assistant promise investment returns. In the end, the agent will refuse to answer questions like “Will I lose money?” or “Can you guarantee profit?” ![Inducing an agent will be blocked by compliance rules specific to certain business scenarios.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-3.png) Inducing an agent will be blocked by compliance rules specific to certain business scenarios. Image by Author ### How will you do it Why use compliance checks as an example? Because in real web apps, product teams don’t manage compliance themselves. The compliance department creates the rules and sends them as microservices to each product. That way, teams don’t touch those rules. They just plug them in using framework middleware. It’s common in normal web apps. We’ll do the same with MAF agents, using middleware to insert compliance logic. To simulate real setups, this project has two parts: one server and one client. ![The compliance check middleware will include both server and client modules.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/01/image-4.png) The compliance check middleware will include both server and client modules. Image by Author On the compliance department side, we’ll deploy a separate agent that reviews messages. It uses an LLM to check user inputs for prompt injections or non-compliant content. On the business side, we’ll have a middleware that intercepts user requests and sends them to that server. It decides whether the agent should respond. The two parts communicate using the AG-UI protocol. ### Server implementation Let’s build the compliance-checking agent server. Since it only checks user requests, I’ll use the `Qwen3-30b-a3b-instruct-2507` model for speed. ```python agent = OpenAILikeChatClient( model_id=Qwen3.Q30B_A3B ).create_agent( name="Assistant", instructions=dedent(""" You are a compliance review officer. You will review user requests or system-generated text for compliance. Your main task is to check user requests and determine whether they are trying to induce the system to produce content that guarantees investment returns or similar topics. You should output a JSON text, like {"is_compliance": 1, "reason": ""} Here, is_compliance being 1 means compliant, and 0 means non-compliant. reason should state the reason for compliance or non-compliance. Only output the JSON text without any markdown formatting, and do not add any introduction or explanation. """), ) ``` To make reviews clear and fast, I ask the agent to output a JSON string. Note: Although I made Microsoft Agent Framework support structured output for Qwen models in a previous article: [Make Microsoft Agent Framework’s Structured Output Work With Qwen and DeepSeek ModelsThings You Always Have to Do When Switching a Framework![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-57.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Agent_Framework_cover-2.webp)](https://www.dataleadsfuture.com/make-microsoft-agent-frameworks-structured-output-work-with-qwen-and-deepseek-models/) I do not know if the issue comes from AG-UI protocol or MAF. **When the agent acts as an AG-UI server, structured output via `response-format` stops working. You must specify the output format in the prompt.** After initializing the agent, use `add_agent_framework_fastapi_endpoint` from `agent_framework_ag_ui` to bind the agent to a FastAPI `app` instance. ```python app = FastAPI(title="AG-UI Server") add_agent_framework_fastapi_endpoint(app, agent, "/compliance") ``` Finally, run it with `uvicorn`: ```python if __name__ == "__main__": import uvicorn uvicorn.run(app, host="0.0.0.0", port=8888) ``` ### Middleware implementation This middleware is more complex, so we’ll use class-based middleware. Here’s the full code: ```python class ComplianceCheckMiddleware(ChatMiddleware): def __init__(self, *args, **kwargs): self._init_compliant_agent() super().__init__(*args, **kwargs) async def process( self, context: ChatContext, next: Callable[[ChatContext], Awaitable[None]], ): check_result: ReviewResults = await self._get_compliance_result(context) if not check_result.is_compliance: self._output_result( context, f"😒We can’t keep providing the service because:\n{fill(check_result.reason)}") return await next(context) @staticmethod def _output_result(context: ChatContext, response: str) -> None: if context.is_streaming: #4 async def output_stream() -> AsyncIterable[AgentRunResponseUpdate]: yield AgentRunResponseUpdate(contents=[TextContent(text=response)]) context.result = output_stream() else: context.result = AgentRunResponse( messages=[ChatMessage(role=Role.ASSISTANT, text=response)] ) async def _get_compliance_result(self, context: ChatContext) -> ReviewResults: messages = [message for message in context.messages if message.role.value == "user"][-5:] response = await self.agent.run(messages) #2 check_result = ReviewResults.model_validate_json(response.text) #3 return check_result def _init_compliant_agent(self) -> None: client = AGUIChatClient( #1 endpoint="http://127.0.0.1:8888/compliance" ) self.agent = client.as_agent( name="compliance_agent", instructions="You’re a compliance officer, and you review user requests." ) ``` A few details to watch: 1. `_init_compliant_agent` creates the AG-UI client but works just like a normal chat client. 2. I sent recent user messages for better review accuracy. **But the `AgentMiddleware` context only holds the latest message. To get the message history, you must use `ChatMiddleware`.** 3. Since AG-UI doesn’t support `response_format`, I parse JSON manually. 4. `_output_result` sends text output if a check fails. It switches based on `context.is_streaming`. Now we can make a business agent. Use a bigger model and a normal system prompt; just remember to load the `ComplianceCheckMiddleware`. ```python chat_client = OpenAILikeChatClient(model_id=Qwen3.NEXT) agent = chat_client.as_agent( name="chat_assistant", instructions="You are a helpful assistant. Answer the user's question in short and simple words.", middleware=[ComplianceCheckMiddleware()] ) ``` Let’s test it with a multi-turn chat client: ```python async def main(): thread = agent.get_new_thread() while True: user_input = input("\nUser: ") if user_input.startswith("exit"): break stream = agent.run_stream(user_input, thread=thread) print("\nAssistant: ") async for event in stream: print(event.text, end="", flush=True) print() ``` You’ll see the agent chats normally most of the time. If you ask about guaranteed returns, it refuses to answer but continues working fine afterward. Task done. --- ## Conclusion In this guide, we explored how middleware works in Microsoft Agent Framework. Middleware lets us add new logic for logging, permissions, or compliance without touching the main agent code or prompt text. In the project section, I used class-based middleware to show how to review user inputs for compliance. We also took a quick look at AG-UI for building agent microservices. This helps when many teams need to make agents collaborate, and I’ll cover AG-UI and A2A in detail later. If you have questions or want to learn more, leave a comment. Don’t forget to subscribe to my blog and share this article with your friends—maybe it’ll help someone build smarter agents 😁. --- Here is the source code of this article: [agentic-ai-playground/11\_MAF\_Middleware\_Basic at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-29.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-27)](https://github.com/qtalen/agentic-ai-playground/tree/main/11%5FMAF%5FMiddleware%5FBasic?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### My Agent System Looks Powerful but Is Just Industrial Trash URL: https://www.dataleadsfuture.com/my-agent-system-looks-powerful-but-is-just-industrial-trash/ Last updated: 2026-04-22T08:20:05.000Z This weekend note is a bit late because Phase One of my Deep Data Analyst project failed for now. That means I can’t continue the promised Data Analyst Agent tutorial. --- ## What Happened? I actually built a single-agent data analysis assistant based on the ReAct pattern. This assistant could take a user’s analysis request, come up with a reasonable hypothesis, run EDA and modeling on the uploaded dataset, give professional business insights and actionable suggestions, and even create charts to back up its points. If you’re curious about how it worked, here’s a screenshot that shows how cool it looked: ![The cool effects of my data analysis agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/data_analysis_report.gif) The cool effects of my data analysis agent. Image by Author After all, this was just a single-agent app. It wasn’t that hard to build. If you remember, I explained how I used a ReAct agent to solve the Advent of Code challenges. Here’s that tutorial: [How I Crushed Advent of Code And Solved Hard Problems Using Autogen Jupyter Executor and Qwen3A detailed guide on using AI to generate code and solve puzzles automatically and quickly![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-52.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/jupyter_executor_cover-1.webp)](https://www.dataleadsfuture.com/how-i-crushed-advent-of-code-and-solved-hard-problems-using-autogen-jupyter-executor-and-qwen3/) If you tweak that agent’s prompt a bit, you can get the same kind of data analysis ability I’m talking about. --- ## Why Do I Call It a Failure? Because my agent, like most that AI hobbyists build, is just one of those: **Perfect for impressing your boss with a beautiful, powerful prototype, but once real users try it, it suddenly breaks down and becomes industrial trash.** --- ## Why Do I Say That? My agent has two serious problems. ### 1\. Very poor robustness This is the top feedback I got after giving it to analyst users. If you try it once, it looks amazing. It uses methods and technical skills beyond a regular analyst to give you a very professional argument. You’d think replacing humans with AI was the smartest move you've ever made. But data analysis is about testing cause and effect over time. You must run the same analysis daily or weekly to see if the assistant’s advice actually works. Even with the same question, the agent changes its hypotheses and analysis methods each run. It then gives different advice each time. That’s what I mean by poor stability and consistency. Imagine you ask it to use an RFM model to segment your users and give marketing suggestions. Before a campaign, it uses features A, B, C and makes five levels for each. After the campaign, it suddenly adds a derived metric D and now segments on A, B, C, D. You couldn’t even run an A/B test properly. ### 2\. It suffers from context position bias If you’ve read my earlier posts, you know my Data Analyst agent runs code through a stateful Jupyter Kernel-based interpreter. [Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent SkillsUse Jupyter code executor to help your agent finish tasks in a smarter way![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-53.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-9.webp)](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) This lets the agent act like a human analyst, first making a hypothesis, running code in a Jupyter notebook to test it, and then coming up with a new hypothesis based on results — iterating over and over. ![Agents based on the ReAct mode will perform EDA like human analysts.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-9.png) Agents based on the ReAct mode will perform EDA like human analysts. Image by Author This gives the agent strong autonomous exploration and error-recovery skills. But here’s the problem. In a past post, I mentioned that LLMs have position bias when dealing with long conversation histories: [Fixing the Agent Handoff Problem in LlamaIndex’s AgentWorkflow SystemThe position bias in LLMs is the root cause of the problem![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-54.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Agents_Hands_Off.webp)](https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/) In short, LLMs don’t treat each message fairly. They don’t weight importance by recency like you think they would. ![LLMs do not assign weights to message history as people might think.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-10.png) LLMs do not assign weights to message history as people might think. Arxiv 2404.01430 As we keep making and testing hypotheses, the history grows. Each message in it matters. The first shows the data structure, a later one proves a hypothesis wrong, so we skip it next time — all important. The LLM doesn’t see it this way. As the process goes, it starts focusing on wrong messages while ignoring the ones that have been fixed. So it repeats mistakes. This either wastes tokens and time or sends the analysis off-track into another topic. Neither is good. So Phase One of my data analysis agent is done. --- ## Any Ways to Fix It? ### Build a multi-agent system with atomic skills For robustness, you’d probably think of using a Context Engineer to lock in the plan and metric definitions before analysis starts. Also, when an analysis works well, we should save the plan and prior assumptions in long-term memory. Both mean giving the agent new skills. But remember, my agent is based on ReAct, which means its prompt is already huge — over a thousand lines now. ![Agents based on the ReAct pattern are often too complex to debug.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-11.png) Agents based on the ReAct pattern are often too complex to debug. Image by Author Adding anything risks breaking this fragile system and disrupting prompt-following. So a single agent won’t cut it. We should split the system into multiple agents with atomic skills, then use some orchestration to bring them together. We can imagine this multi-agent app as a coordinate system with at least these agents: - **Issue Clarification Agent** — asks the user questions to clarify the problem, confirm metrics, and scope. - **Retrieval Agent** — pulls metric definitions and calculation methods from a knowledge base, plus analysis methods written by real data scientists. - **Planner Agent** — proposes prior hypotheses, sets an analysis approach, and makes a full plan to keep later agents on track. - **Analyst Agent** — breaks the plan into steps, uses Python to execute them, and tests the prior hypotheses. - **Storyteller Agent** — turns complex technical results into engaging business stories and actionable advice for decision-makers. - **Validator Agent** — ensures the whole process is correct, reliable, and business-compliant. - **Orchestrator Agent** — manages all the agents and assigns tasks. ![My new design for the multi-agent data analyst.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-12.png) My new design for the multi-agent data analyst. Image by Author ### Choose the right agent framework We need an agent framework that supports message passing. When a new task comes up or an agent finishes, a message should go to the orchestrator. The orchestrator should also send tasks by message. The framework should support context state saving. Agents’ intermediate results should go to the context, not all to the LLM, so position bias doesn’t get in the way. If you ask GPT, it will recommend LangGraph and Autogen. I’d skip LangGraph. Even though its workflow is fine, its agents still run on LangChain, which I just don’t like. When people compare Autogen with others, they say Autogen is better for research-heavy tasks like data analysis that need more autonomy. But Autogen’s Selector Group Chat, while good for orchestrators, can’t manage message history well. You can’t control what goes to the LLM, and orchestration is a black box. Autogen’s GraphFlow is also half-baked. Workflow only supports agent nodes and no context state management. [I Used Autogen GraphFlow and Qwen3 Coder to Solve Math Problems — And It WorkedMore reliable than your math professor![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-55.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-10.webp)](https://www.dataleadsfuture.com/i-used-autogen-graphflow-and-qwen3-coder-to-solve-math-problems-and-it-worked/) The bigger risk: Autogen has stopped development. For a 50k-star agent framework, that’s a shame. ### What about Microsoft Agent Framework (MAF)? I like it. Easy to use, takes good ideas from earlier frameworks, and avoids their mistakes. I’m ready to use it with Qwen3 and DeepSeek: [Make Microsoft Agent Framework’s Structured Output Work With Qwen and DeepSeek ModelsThings You Always Have to Do When Switching a Framework![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-56.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Agent_Framework_cover-1.webp)](https://www.dataleadsfuture.com/make-microsoft-agent-frameworks-structured-output-work-with-qwen-and-deepseek-models/) I’m studying MAF’s Workflow feature now. It’s nice: multiple node types, context state management, OpenTelemetry observability, and orchestration modes like Switch-Case and Multi-Selection. It has almost everything I want. It also feels ambitious. With new abilities like MCP, A2A, AG-UI, and Microsoft backing it, MAF should have a better long-term future than Autogen. --- ## My Next Steps I’m reading MAF’s user guide and source now. I’ll start using it in my agent system. I’m still working on Deep Data Analyst. After switching frameworks, I’ll need to adapt things for a while. The good news: a multi-agent system lets me add skills step by step, so I can share and show progress anytime instead of waiting until the whole project is done. 😂 I also want to explore Workflow’s potential in MAF. I’ll see if it can handle different AI agent design patterns. That will help us understand how to use this promising framework. What are you interested in? Leave me a comment. Don’t forget to subscribe to my newsletter to get my latest agent research in your inbox without waiting. And share my blog with your friends — maybe it can help more people. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Make Microsoft Agent Framework’s Structured Output Work With Qwen and DeepSeek Models URL: https://www.dataleadsfuture.com/make-microsoft-agent-frameworks-structured-output-work-with-qwen-and-deepseek-models/ Last updated: 2026-04-22T08:20:29.000Z **Update:** After version `python-1.0.0b260114`, MAF made big changes to the `response_format` related APIs, but the official docs haven’t been updated yet. This article is based on version `python-1.0.0b260127` and brings you the latest structured output solution. ## Introduction Today, we’ll add some extra features to the Microsoft Agent Framework so that Qwen and DeepSeek can also utilize structured output. The main reason is that Autogen has stayed on version v0.75 for a long time, which makes it necessary to switch to Microsoft Agent Framework soon. Every time we switch the agent framework, we have to make it work with some common LLMs. This time is no exception. Luckily, Microsoft Agent Framework is pretty easy to use. We just need to adapt the structured output feature, and we can use it right away. As usual, I’ll put the source code at the end of the article for you to use. --- ## Background On Structured Output ### How does Agent Framework do structured output? In Microsoft Agent Framework, we set the `response_format` parameter to a Pydantic `BaseModel` data class to tell the LLM to produce structured output, like this: ```python from pydantic import BaseModel class PersonInfo(BaseModel): """Information about a person.""" name: str | None = None age: int | None = None occupation: str | None = None response = await agent.run( "Please provide information about John Smith, who is a 35-year-old software engineer.", options={ "response_format": PersonInfo }, ) ``` There are two places to set the `response_format` parameter: 1. Set it during the `ChatAgent` initialization in `default_options` parameter. This becomes a global parameter for the agent, and all later communications with OpenAI-compatible models use it. 2. Set it when calling `run` or `run_stream` in `options` parameter. This works only for that single API call. The `response_format` set in `run` or `run_stream` is higher priority than the setting in the `ChatAgent` creation. That means the `response_format` in `run` will override what was set when creating the `ChatAgent`. ![The conversion process of the response_format parameter in Microsoft Agent Framework. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-4.png) The conversion process of the response\_format parameter in Microsoft Agent Framework. Image by Author By default, we use `OpenAIChatClient` to call OpenAI’s API. Before the API call, a `_prepare_options` method converts the `BaseModel` into `{"type": "json_schema", "json_schema": }` and passes it to the LLM. So that’s how Agent Framework makes the LLM do structured output. Our extension will go into the `_prepare_options` method of `OpenAIChatClient`. ### Do Qwen and DeepSeek support json\_schema settings? According to the official docs, both Qwen and DeepSeek support structured output. But they only support setting the OpenAI client’s `response_format` to `{"type": "json_object"}` and require the keyword `json` in the prompt to enable structured output. They do not support OpenAI’s API way of setting `response_format` to `json_schema`. If we don’t extend the Microsoft Agent Framework and force `response_format` to be a `BaseModel` class, we’ll see errors like this: ```text Error code: 400 - {'error': {'message': "<400> InternalError.Algo.InvalidParameter: 'messages' must contain the word 'json' in some form, to use 'response_format' of type 'json_object'.", 'type': 'invalid_request_error', 'param': None, 'code': 'invalid_parameter_error'}} ``` So for Qwen and DeepSeek, without modifying the Microsoft Agent Framework, we can’t use the structured output feature. ### How to make Qwen and DeepSeek output using json\_schema Even though Qwen and DeepSeek don’t support `{"type": "json_schema"}`, we can still inject `json_schema` into the system prompt so the LLM outputs according to our data class. The trick is: before calling the OpenAI API, convert the `BaseModel` to its `json_schema`, attach it to the system prompt, and send it along. If you want to know exactly how I made Qwen output according to a Pydantic BaseModel’s rules, read my popular article where I explain multiple methods for this: [Build AutoGen Agents with Qwen3: Structured Output & Thinking ModeSave yourself 40 hours of trial and error![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-49.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/structured_output-1.webp)](https://www.dataleadsfuture.com/build-autogen-agents-with-qwen3-structured-output-thinking-mode/) --- ## How I Extended It Now, let’s see exactly how to extend Microsoft Agent Framework so Qwen and DeepSeek can do structured output. I know you want the answer fast, so here’s the modified code you can use right now: ```python from typing import override, MutableSequence, Any from textwrap import dedent from copy import deepcopy from pydantic import BaseModel from agent_framework.openai import OpenAIChatClient from agent_framework import ChatMessage, ChatOptions, TextContent class OpenAILikeChatClient(OpenAIChatClient): @override def _prepare_options(self, messages: MutableSequence[ChatMessage], options: dict[str, Any]) -> dict[str, Any]: chat_options_copy = deepcopy(options) response_format = chat_options_copy.get("response_format") if ( response_format and isinstance(response_format, type) and issubclass(response_format, BaseModel) ): structured_output_prompt = self._build_structured_prompt(response_format) if old_instructions := chat_options_copy.get("instructions"): chat_options_copy["instructions"] = f"{old_instructions}\n\n{structured_output_prompt}" else: messages = [ChatMessage(role=Role.SYSTEM, text=structured_output_prompt), *messages] chat_options_copy["response_format"] = {"type": "json_object"} return super()._prepare_options(messages, chat_options_copy) @staticmethod def _build_structured_prompt(response_format: type[BaseModel]) -> str: json_schema = response_format.model_json_schema() structured_output_prompt = dedent(f""" \n Your output must adhere to the following JSON schema format, without any Markdown syntax, and without any preface or explanation:\n {json_schema}\n """) return structured_output_prompt ``` As I said before, both `run` and `run_stream` call `OpenAIChatClient`’s `_prepare_options` method, so it’s the best place to extend. I marked each part of the code with numbers in the comments so I can explain in order: 1. The `chat_options` object is the parameters you pass to the method. We need to `deepcopy` it to a new object because we’re going to change `response_format` to `{"type": "json_object"}` to work with DeepSeek. Agent Framework still needs the original `BaseModel` to convert the returned JSON string back to a data class. 2. Then we take the `json_schema` from the `BaseModel`, turn it into part of the system prompt, and wrap it with `xml` tags. 3. The original `_prepare_options` checks if `messages` is empty. We’ll only handle the case where `messages` is not empty, meaning the user sends at least a user message. 4. If the first message in `messages` is a system message, we attach the structured output prompt to the system message, replacing the old system message. 5. If the first message is a user message, we create a new system message with just the structured output prompt and put it at the front of the `messages` list. With this change, Microsoft Agent Framework now supports structured output for Qwen and DeepSeek. Next, let’s test some common cases to make sure it works. --- ## Testing the Extension ### Prepare an MLflow server to observe Before testing, we need a monitoring tool to check the messages Agent Framework sends to the LLM API. Agent Framework supports logging platforms based on `opentelemetry`, but it doesn’t log system messages by default, so that won’t work for our case today. ![Agent Framework's OpenTelemetry output doesn't log the system message used when calling the LLM.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-5.png) Agent Framework's OpenTelemetry output doesn't log the system message used when calling the LLM. Image by Author In a previous article, I showed how I use MLflow to see the messages sent to OpenAI’s API: [Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracing TutorialEnhance your multi-agent application’s observability, explainability and Traceability![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-50.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Tracing-with-MLflow-3.x-1.webp)](https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/) So today we’ll still use MLflow’s `openai.autolog` API, because it can record system messages sent to the LLM. You just need to start a `server` like this: ```shell mlflow server --host 0.0.0.0 --port 5000 ``` Then in the test code, add a call to `openai.autolog`: ```python mlflow.set_tracking_uri(os.environ.get("MLFLOW_TRACKING_URI")) mlflow.set_experiment("Default") mlflow.openai.autolog() ``` ### Test single-turn conversation First, let’s follow the official docs to test normal structured output. Set up a data class, then set it in the `run` method: ```python class PersonInfo(BaseModel): """Information about a person.""" name: str | None = None age: int | None = None occupation: str | None = None async def main(): response = await agent.run( "Please provide information about John Smith, who is a 35-year-old software engineer.", options={ "response_format": PersonInfo }, ) print(response.text) ``` Check on MLflow: ![The json_schema prompt has already been appended to the system prompt.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-6.png) The json\_schema prompt has already been appended to the system prompt. Image by Author We can see the data class has been turned into a `json_schema` prompt, attached to the system prompt. Also, we can get the structured object directly through `response.value`. ### Test multi-turn conversation Now let’s test Microsoft Agent Framework’s multi-turn example. First, set a `response_format` at `create_agent`, without setting it in `run`: ```python class OutText(BaseModel): output: str agent = client.create_agent( instructions="You are a good assistant.", name="assistant", default_options={ "response_format": OutText }, ) async def main(): result1 = await agent.run( "How many kilometers is the highway from Wuhan to Beijing?", thread=thread, ) print(result1.text) ``` Then use `run_stream` for the second turn and set another `response_format`: ```python class ETA(BaseModel): hours: int final_response = await AgentRunResponse.from_agent_response_generator( agent.run_stream( "How long would it take to drive there at 120 km/h?", thread=thread, options={ "response_format": ETA }, ), output_format_type=ETA ) print(final_response.value) ``` Check on MLflow: ![The first round of conversation used the default response_format parameter.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-7.png) The first round of conversation used the default response\_format parameter. Image by Author ![The second round of conversation switched to the response_format parameter passed into the run_stream method.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-8.png) The second round of conversation switched to the response\_format parameter passed into the run\_stream method. Image by Author No problems at all. The `response_format` in `run_stream` overrides the one set in `create_agent` as expected. --- ## Conclusion With Autogen no longer updated, we’ve started moving to Microsoft Agent Framework. During this migration, we extended the Microsoft Agent Framework so Qwen and DeepSeek can use structured output. I hope Qwen and DeepSeek’s APIs will one day support setting `response_format` to `{"type": "json_schema"}` directly, so we wouldn’t have to adapt the framework every time we switch. Structured output is just about adding a `json_schema` description in the system prompt so the LLM outputs content as we define. So even if you’re not using Microsoft Agent Framework, you can modify things in a similar way. That’s it for today’s journey. If you find this tutorial useful, please share it with your friends. --- Here’s the source code for today’s tutorial: [agentic-ai-playground/10\_Agent\_Framework\_Qwen3\_DeepSeek at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-22.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-20)](https://github.com/qtalen/agentic-ai-playground/tree/main/10%5FAgent%5FFramework%5FQwen3%5FDeepSeek?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### A Quick Guide to Containerizing Agent Applications with Podman URL: https://www.dataleadsfuture.com/a-quick-guide-to-containerizing-agent-applications-with-podman/ Last updated: 2026-04-22T08:21:03.000Z ## Introduction For enterprise-level agent applications, the best way to safely run code generated by agents is to use containerization. This isolates the code execution environment from your server’s operating system. In a previous article, we built a code interpreter sandbox based on a Jupyter container. We proved that once an agent has access to a stateful code runtime, it gains the ability to solve complex problems and perform data analysis: [Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent SkillsUse Jupyter code executor to help your agent finish tasks in a smarter way![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-47.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-7.webp)](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) However, Docker Desktop is off-limits in most enterprises due to its commercial license restrictions. Yet our multi-agent development work absolutely depends on a containerized environment. So we must find a suitable alternative. Ideally, one fully compatible with Docker, so it works seamlessly with existing agent frameworks that rely on the Docker client. Podman is exactly what we need. Developed by Red Hat, it’s an open-source container management tool that runs on both Mac and Windows. It can fully replace Docker Desktop in your development setup. This short post will help you quickly set up Podman so you can start building agent applications with code interpreters right away. If you’d like deeper background knowledge and a more thorough introduction to Podman, I highly recommend the [**Podman for the Absolute Beginners - Hands-On DevOps course on Udemy**](https://www.udemy.com/course/podman-for-the-absolute-beginners-hands-on-devops/?couponCode=KEEPLEARNING&ref=dataleadsfuture.com)**.** --- ## Quick Start ### Prerequisites This guide assumes you’re developing on Windows 11\. From what I know, setup on Mac is even simpler. Podman runs its host inside a Linux system on WSL2\. So before proceeding, make sure you’ve already installed WSL2 on your machine. If your company requires a VPN to access the internet, you’ll also need to configure WSL2\. To do this, create a `.wslconfig` file in your Windows `%USERPROFILE%` directory (your user folder) with the following content: ```latex [experimental] autoMemoryReclaim=gradual networkingMode=mirrored dnsTunneling=true firewall=true autoProxy=true ``` ### Install Podman As mentioned, Podman’s host runs in a WSL2-based Linux environment called `podman-machine-default`. So when installing Podman, you actually need to set up this `podman-machine-default`. One approach is to install Podman Desktop first, then create the `podman-machine-default` through its interface. But for some reason, this always failed for me—the installer would just hang. So I went the other way: I installed `podman-machine` standalone first, then installed Podman Desktop. Go to Podman’s [GitHub Releases page](https://github.com/containers/podman/releases?ref=dataleadsfuture.com), download the latest `.msi` installer, and run it. ![Find and download the .msi installer under the latest release.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image.png) Find and download the .msi installer under the latest release. Image by Author After installation, create your `podman-machine-default` host: ```shell podman machine init --rootful=true ``` Important note: if your agent framework uses the Docker SDK to launch containers, you **must** set `--rootful=true`. I’ll explain why later. Next, install Podman Desktop from [here](https://podman-desktop.io/?ref=dataleadsfuture.com). Just click “Next” all the way through. Once done, open Podman Desktop from your taskbar, go to the **Settings** tab, and check **Resources**. You should see the `podman-machine` you just created. That means Podman Desktop is ready. ![If you see this screen, it means your Podman Desktop is already installed.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-1.png) If you see this screen, it means your Podman Desktop is already installed. Image by Author But since your agents will use the Docker SDK to manage containers, you need to enable Docker compatibility in Podman Desktop, as shown below: ![We use the Docker SDK to manage containers in our agent framework, so make sure Docker Compatibility is turned on.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-2.png) We use the Docker SDK to manage containers in our agent framework, so make sure Docker Compatibility is turned on. Image by Author This setting also enables Podman Compose, which lets you route `docker compose` commands to `podman compose`. ### Set up proxy configuration If your system uses a system-wide proxy, Podman Desktop will automatically pick it up. But here’s the catch: Podman Desktop applies these settings inside the `podman-machine` Linux system. Unlike Windows, where bypass domains are separated by semicolons, Linux uses commas. You’ll need to manually adjust this, or your `podman-machine` might lose external network access. ![You should use the Linux path separator here.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/12/image-3.png) You should use the Linux path separator here. Image by Author ### Create a soft link for the storage path By default, Podman stores data in `%USERPROFILE%\.local\share\containers\podman`. I really dislike keeping data files on the C drive. I prefer moving them to D drive—for example, `D:\Documents\AppData\podman`. Here’s how: copy everything from the C-drive `podman` folder to your new D-drive location, delete the original C-drive folder, then create a junction link (remember to replace `%USERPROFILE%` with your actual path): ```shell mklink /j C:\Users\qianpeng\.local\share\containers\podman D:\Documents\AppData\podman ``` Now everything works normally, but your data lives safely on D drive—no risk of losing it during a system reinstall. ### Test your Podman installation If you followed my [earlier article](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) and created a `Dockerfile` for a `jupyter-server`, you can now test whether Podman builds images correctly. Navigate to your `Dockerfile` directory and run: ```shell podman build --network=host -t jupyter-server . ``` Remember, the `podman-machine` host is network-isolated from your Windows system. The `--network=host` flag lets the build process access your Windows network, making it easier to pull base images and install Python packages from PyPI. Note: this flag only affects image building—it has no impact when running containers. ### Make Docker CLI commands work If you’re used to typing `docker` commands, here’s a neat trick: create a `docker.bat` file in `%USERPROFILE%/.local/bin` (Make sure this path is in your system `PATH`). Put this inside `docker.bat`: ```shell @echo off setlocal EnableDelayedExpansion if "%~1"=="build" ( shift rem Default to --network=host podman build --network=host %2 %3 %4 %5 %6 %7 %8 %9 ) else ( podman %* ) endlocal ``` Now you can keep using `docker` commands as usual. Since Podman’s CLI is mostly compatible with Docker’s, everything feels the same—except when you run `docker build`, it automatically adds `--network=host` to the underlying `podman build` command. ### Call Docker SDK from inside a containerized app If your agent app runs inside a container and needs to use the Docker SDK to spin up Python code interpreter containers (a common pattern across agent frameworks), you’ll need to let the SDK talk to Podman. Do this by mounting the `podman.sock` socket into your app container so the Docker SDK can reach it: ```shell docker run -v /run/podman/podman.sock:/var/run/docker.sock --rm app-test ``` Again, your `podman-machine-default` **must** be initialized with `--rootful=true`. ### Fix Docker SDK timeout issues When your agent generates code and sends it via Docker SDK to a code interpreter container, there’s often a delay between code generations while the LLM produces tokens. On the second SDK call, you might hit a timeout like this: ```latex requests.exceptions.ConnectionError: ('Connection aborted.', RemoteDisconnected('Remote end closed connection without response')) ``` This happens because `podman-machine` sets a very short `service_timeout` for its engine by default. Fix it by SSHing into your `podman-machine`: ```shell podman machine ssh ``` Now you’re in the Linux VM. Edit `/etc/containers/containers.conf`: ```shell sudo vi /etc/containers/containers.conf ``` Set `service_timeout` to 0 (meaning “never time out”) under the `[engine]` section: ```shell [engine] cgroup_manager = "cgroupfs" service_timeout = 0 ``` Since this is a dev environment, security isn’t a concern here. ### Adjust Podman container stop timeout During development, you often start and stop your agent program from the command line. But you might notice it takes forever to exit after the program finishes. That’s because Podman waits by default (10 seconds!) for containers to gracefully shut down after receiving a stop signal. Shorten this by editing the same config file: ```shell [engine] stop_timeout = 2 ``` Now containers stop in 2 seconds—much snappier! ### Allow Containers to Use the --restart Flag Podman lets you use the `--restart` flag when starting a container, just like Docker, so your containers can automatically start up again after your computer reboots. The only thing you need to do first is set up `podman-restart.service` inside your podman-machine. Here's how: First, jump into podman-machine-default and run these shell commands: ```shell # Check the status of podman-restart.service systemctl status podman-restart.service # Enable the service (so it starts on boot) sudo systemctl enable podman-restart.service # Start the service right now sudo systemctl start podman-restart.service ``` Then create your container with the `--restart` flag: ```shell podman run -d -p 5000:5000 --name mlflow-server --restart=always mlflow-server ``` Or if the container already exists, just update its restart policy: ```shell podman update --restart=always mlflow-server ``` Finally, to make sure everything is set up correctly, run this in your Windows terminal: ```shell podman inspect mlflow-server --format='{{.HostConfig.RestartPolicy.Name}}' ``` And that's it, you're good to go! --- ## Conclusion Ever since Gemini 3.0 accidentally deleted 800GB of user data, running agent-generated code inside containers has become essential. And yes—stateful code interpreters truly boost an agent’s ability to tackle complex problems. Most agent frameworks rely on the Docker SDK to manage containers. But Docker Desktop’s licensing blocks its use in enterprise dev environments. This guide shows you how to use Podman Desktop as a drop-in replacement. Following these steps, you can quickly build a Docker-compatible container dev environment. I’ve used this exact setup for over six months with zero issues. This article covers only the minimal setup needed for agent development. If you want to dive deeper into Podman’s internals and advanced usage, check out the official tutorials. Our journey toward **Deep Data Analyst Agents** continues—stay tuned for the next episode! ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Share My LLM Prompts and Tips That Make Work and Learning Super Efficient URL: https://www.dataleadsfuture.com/share-my-llm-prompts-and-tips-that-make-work-and-learning-super-efficient/ Last updated: 2026-07-24T01:54:14.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* A lot of friends ask me how I manage to stay busy with work every day, yet still find time to learn and write blog posts. The answer is simple: I use AI to help me learn about AI. Today I’m sharing all the AI tools, tricks, and prompts I’ve used over the past two years at work. No fluff—just straight-up useful stuff. --- ## Find a Good AI Client If you use LLMs to boost your daily productivity, chatting with the model is still the main way most people interact with it. That means having a solid AI client app is essential. My favorite AI client right now is [Cherry Studio Community Edition](https://www.cherry-ai.com/?ref=dataleadsfuture.com). It supports multiple languages, is completely open source and free, lets you connect to all kinds of model services, and even lets you add your own System Prompt and MCP tools. These features form the foundation for all the tips I’ll share next. You can pick any model service you like. I recommend [OpenRouter](https://openrouter.ai/models?ref=dataleadsfuture.com)—with one API key, you get access to Gemini, GPT, Claude, Qwen, and many other commercial or open-source models. ![Configure the model service in Cherry Studio's settings interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-19.png) Configure the model service in Cherry Studio's settings interface. Image by Authro Then there’s the agent interface. You can fill in different System Prompts based on your use case. The prompts I use daily go right here. ![Enter the prompt you want to share today in the agent interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-21.png) Enter the prompt you want to share today in the agent interface. Image by Author Finally, in the chat interface, you can add the agent you just set up as your assistant, tweak your model settings, and start chatting with the LLM. ![Just add the agent you just set up as an assistant, and you're good to go!](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-23.png) Just add the agent you just set up as an assistant, and you're good to go! Image by Author Best of all, using an AI desktop client lets you keep your knowledge base local and connect to self-hosted LLM services, giving you and your company the strongest data security possible. --- ## Some Suggestions on Models and Settings Next up are the models and parameters I use daily. These aren’t “correct” answers—just my personal experience. ### Model choices I follow a simple rule: if I’m using it myself, I go with the best. So I stick to commercial models unless I’m building agents, where I might pick an open-source model based on need. Here’s what I use regularly: **Gemini 3.0 Pro** is my top pick for vibe coding. Right now, code generated by Gemini 3.0 has the highest accuracy, which saves me tons of debugging time. **GPT-5** is the classic reliable choice—a great balance between cost and expertise. I use it for everything except coding. **Qwen3 Max**… well, I really dislike its overly encouraging tone. It always tells me I’m doing great, no matter what, and that drives me nuts. But I have to admit—Qwen3 shines in localization and language handling. I use it whenever I need translation or proofreading. ### Parameter settings If you’ve read my articles before, you probably already know what each LLM parameter does. Here’s how I personally set them when using models myself. - **Temperature.** Lower values make the model more predictable; higher values make it more creative. I adjust based on context. For coding, I set `temperature` to 0.01 to keep responses consistent across chats. For everything else, I stick with the default 0.7—it feels like talking to a real person. Later, I’ll show you how to make the LLM write creative motivational text—I crank `temperature` up to 0.8–0.9 for that. - **Context length.** Few people pay attention to this setting, but besides saving token costs, it can offer unexpected benefits. For my translation agent, I set context length to 2—meaning no chat history is kept. That way, the LLM only translates what I input right now, without interference from past translations. - **Max tokens.** This controls the maximum number of tokens per response. I always adjust it. For reasoning tasks, I keep it low—otherwise wait times and token costs explode. For writing, I set it high to avoid cutting off long articles due to default limits. ### Try using MCP I’ve always felt that MCP was built exactly for making LLM clients more powerful. With MCP, your personal chat interface can unlock all kinds of agent capabilities. Here’s an example: We all know that due to design choices and legal restrictions, the built-in web search in LLMs keeps getting worse. Default search rarely gives useful info, especially over multi-turn conversations. But you can install a `tavily-mcp` service for better web search. Just sign up on their site, get an API key, and add the tavily-mcp JSON config to your client. ![You can find the JSON for configuring MCP on the official websites of the various tools.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-24.png) You can find the JSON for configuring MCP on the official websites of the various tools. Image by Author Traditional search works like this: take your keywords, search the web, then build an answer from the results. MCP-based search is different. During the conversation, the LLM dynamically decides if it needs to search—and generates its own search keywords based on context. This leads to much more accurate results. ![When using tavily-mcp, the agent will try searching with multiple keywords until it finds the answer.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-25.png) When using tavily-mcp, the agent will try searching with multiple keywords until it finds the answer. Image by Author Other MCP tools are great too. `fetch` grabs webpage content from any URL you give it. `memory` uses a knowledge graph to remember key info from your chats, helping you build your own custom AI agent. Now that we’ve covered LLM setup tips, let’s move on to my prompt-writing tricks. If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. --- ## My Prompt-Writing Tips ### Is structured JSON prompting really necessary? After Gemini 3.0 launched, people noticed that using JSON-formatted prompts seemed to help LLMs follow instructions better. That sparked debate: should we always use JSON for clearer, more precise structure? I’ve discussed this several times with the brilliant engineers at Qwen. Their answer was clear: It depends on what format of the training data used during model training. LLMs essentially memorize knowledge—including input formats—through their parameters. If most training text were in Markdown, then Markdown is naturally the best fit. That’s why LLMs output in Markdown by default—they were trained on Markdown-heavy data. So their most familiar format is still Markdown. ![What you think is a clearer format doesn't mean the LLM sees it that way too.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-26.png) What you think is a clearer format doesn't mean the LLM sees it that way too. Image by Author It’s like someone who eats bread every day telling a rice-eating kid that bread is the real staple food—completely forgetting that for the kid, rice is the staple. So my conclusion? Just stick with Markdown—it’s plenty structured. And with modern LLMs, even plain-text instructions work fine as long as you’re clear. ### A universal template for system prompts Just like I usually write articles in a three-part structure (introduction, body, conclusion), having a prompt template saves you from staring at a blank screen. What template works for system prompts? I use the “Who, Can, Do” framework. Every good prompt includes: **role**, **what to (not) do**, and **how to do it**. I define each part with a subheading. Here’s an example: ```text ## Role You are a data analyst skilled at breaking complex tasks into Python-solvable subtasks. ## Tasks 1. **Task breakdown**: Split the user request into substeps, each solvable with Python. 2. **Code generation**: Turn the current substep into Python code. 3. **Code execution**: Run the code using a tool and get the result. 4. **Iterate**: Use the result to decide the next step. Repeat steps 1–3 until you have a final answer. 5. **Insight & advice**: Add thoughtful, practical insights based on the analysis. ## Requirements - Execute one step at a time. No skipping or combining. ## Output - Use Markdown with a clear structure. - Keep tone friendly but authoritative. - Add emojis for warmth. - Format numbers with commas (e.g., 1,000). ``` Here’s what each section means: - **Role** tells the LLM “who I am and what I can do.” “Who I am” shapes output style—serious or playful—based on the role you assign. “What I can do” sets initial boundaries. For MoE models, it can even influence which expert module activates. - **Tasks** are your specific instructions. Since I recommend Markdown, use ordered lists if steps must run in sequence; otherwise, use unordered lists. Lists make your intent crystal clear. - **Requirements** remind the LLM of its limits. Model makers train their LLMs to answer everything—but reality isn’t like that. Explicitly stating what it *can’t* do reduces hallucinations. - **Output** guides output format and tone. Use this section when you care about style, structure, or voice. You can add extra sections based on your needs. For a coding agent, add a “Code Style” section: ```text ## Code Style - Code runs in Jupyter. Reuse existing variables. - Write incrementally and leverage kernel state to avoid repetition. ``` If you use RAG or want to show the LLM how to think, add an “Examples” section: ```text ## Examples When writing Python, wrap code in markdown Python blocks: ```python x=3 ``` You can reuse the variable later: ```python print(x) ``` ``` ### Let the LLM help you debug your prompt Tweaking prompts is expensive—especially when a perfectly tuned prompt stops working after switching models. What can you do besides starting over? Use the LLM itself! After setting your System Prompt, just ask: “**Please repeat in detail what you understand my instructions to be.**” It’s like asking your friends to repeat back a task before they start—to confirm understanding. ![Ask the LLM to repeat the instructions I gave it.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-18.png) Ask the LLM to repeat the instructions I gave it. Image by Author This makes the LLM instantly return its interpretation of your system prompt. Compare it to your original intent and spot gaps. Or go further: ask “**Please use Markdown to repeat in detail what you understand my instructions to be.**” The LLM will then write a clean Markdown version of its understanding. You can copy the good parts straight back into your own prompt. Trust me—it will follow what it writes. I’ve tested this countless times. Beyond theory, I’ll now share some of my go-to prompt examples. You can use them directly at work or as inspiration for your own prompts. Let’s dive in. --- ## My Work Prompt Examples ### Prompt for blog cover images Let’s start with the prompt I use to generate blog cover art. If you’ve read my posts before, you’ll notice a consistent style: a cute little rabbit busy doing various things. I had DeepSeek generate the image prompt, then used DALL·E 3 to create the picture. Even though OpenAI says DALL·E 3 boosts prompts automatically, I still get better results by first using an LLM to write a full image prompt. Here’s what I give DeepSeek: ```text ## Role You are a visual artist skilled at writing DALL·E 3-friendly prompts. ## Task Rewrite my [scene description] into a detailed English prompt perfect for DALL·E 3. ## Length Describe in 5 bullet points. Only the prompt—no intro or explanation. ## Style Colorful illustration on slightly yellowed parchment paper, filled with tech elements. ## Visuals Include impressive details like camera angle and lighting. ``` I used this exact prompt in a previous post about generating ink-wash style illustrations: [Use LLamaIndex Workflow to Create an Ink Painting Style Image Generation WorkflowAdd strong artistic flair through fine control of LLM context![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-46.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/generated_00-1.webp)](https://www.dataleadsfuture.com/use-llamaindex-workflow-to-create-an-ink-painting-style-image-generation-workflow/) Even though that post automated the whole workflow, you can still manually generate the image prompt first, then create the picture. ### Daily translation assistant Since I started blogging, I have often chatted with readers from around the world and answered their questions. I want my English to sound natural and conversational—not stiff like machine translation. So I use an LLM with this prompt: ```text ## Role You are an expert Chinese-English translator in computer science and programming. ## Task - Detect the language of the user’s message. - If it’s not Chinese, translate it into Chinese. - If it’s Chinese, translate it to English. ## Style - Keep translations simple and clear. Avoid complex words. - Use vocabulary a middle schooler would understand. - Sound conversational—like a good friend chatting with you. ## Rules 1. Only translate—never do anything else. 2. Output only the translation—no intro or notes. ## Special Terms Translate these terms as follows: [Chinese phrase]: [English translation] 大模型: LLM 大语言模型: LLM 私有化部署: self-hosted ----------------------------------------------- Now translate this: ``` This prompt automatically detects the input language and translates it into the corresponding language: ![Automatically detect the language and translate it. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-27.png) Automatically detect the language and translate it. Image by Author A few notes: 1. I set context length to 2—only the current text and translation are kept. This prevents interference from past chats. 2. I lock down translations for special terms because models often disagree on them. 3. I end with “Now translate this:”. Even though LLMs use Transformer architecture, they still predict the next token based on prior ones. Sometimes the input is a reader’s question—if I don’t add this line, the LLM answers the question instead of translating it. This phrase makes the task clear. ### Article translation prompt English isn’t my first language, so I always need to translate my new articles. At first, I used DeepL plus Grammarly for polishing. Paying for two subscriptions every month got expensive—and let’s be honest, machine translation isn’t great. So as soon as GPT-3.5 came out, I switched to using LLMs for translation. Even the article you’re reading now was translated by an LLM: ```text ## Role You are a senior data science expert and blog editor. ## Task I wrote a Chinese data science article. Translate it into English. ## Rules - Keep my original paragraph breaks. - Never add code that wasn’t there. - Avoid adverbs, prepositions, and passive voice. - **Only translate—don’t rewrite or change my content.** ## Style - Conversational, light, and cheerful. - Don’t bold the first word or phrase in list items. - Use Title Case for headings and Sentence case for subheadings. ## Tone - Keep it simple and easy to understand. - Use words any US 12th grader would know. - Sound like you’re chatting with a good friend. ## Audience Beginners in data science and people curious about the field. ## Special Terms Translate these terms as follows: [Chinese phrase]: [English translation] 大模型: LLM 大语言模型: LLM 私有化部署: self-hosted ``` ### Data science research assistant My day job is data science, so I trained a specialized assistant—not a general-purpose one. I want to write code that matches my habits. Over time, these preferences settled into the “Requirements” section: ```text ## Role You are a data scientist acting as my programming assistant. Help me improve my data science coding and algorithm skills. ## Requirements 1. Be truthful and precise. Never make things up. 2. Ensure all code runs correctly. 3. Use Python 3.12+ features, syntax, APIs, and best practices. 4. Always use the latest versions and APIs of third-party libraries. 5. Write clean, efficient, readable code. 6. Comment only when necessary. 7. Use vertical bar | for type unions. 8. Prefer `with` statements. 9. Use `pathlib` for file and directory operations. ## Tool Use * tavily-search: Always use `"search_depth": 'advanced'`. ## Response Be truthful, thorough, well-organized, and accurate. ``` Remember the MCP tip earlier? Here I tell the LLM to use tavily’s `advanced` mode for deep web searches when it lacks info. Honestly, with this setup plus Gemini 3.0 Pro, the LLM has massively boosted my data science learning—high efficiency and minimal hallucinations. ### General-purpose daily assistant This one’s simple. I just want honest, helpful answers—no fake positivity like Qwen3-Max’s constant praise. It’s great for everyday questions, and Qwen3-Max is cheaper too: ```text ## Role You are my personal assistant. Give me sincere, useful advice for life and work. ## Requirements * All knowledge, sources, and text must be real and accurate. No fabrications. * Never lie. If you don’t know something, say so. * Don’t just agree with me—point out problems or suggest improvements. ## Output * Never use em dashes in your replies. ``` --- ## Wrap-Up That’s all the LLM tips and prompt examples I’ve gathered over the past two years. They’re not perfect, but they’ve seriously boosted my daily productivity. I hope they help you, too. I’ll keep updating this post with more practical tricks. If there’s something you’d like to see, leave me a comment! ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM AI Developer Professional Certificate**](https://imp.i384100.net/B5eg04?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. ### Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent Skills URL: https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/ Last updated: 2026-04-22T08:22:35.000Z This tutorial will use a more general approach to fully recreate the core tech behind code interpreter sandboxes in commercial products like Manus and Claude agent skills. As always, the source code is at the end of this post. Feel free to grab it. --- ## Introduction Recently, an incident with Gemini 3.0 generated code wiped out 800GB of customer data. This again reminds us of the importance of building code sandboxes for agents. In earlier articles, we’ve already seen many times that letting agents generate code and run it in a sandbox environment can boost their math problem-solving ability and help them tackle complex challenges. [I Used Autogen GraphFlow and Qwen3 Coder to Solve Math Problems — And It WorkedMore reliable than your math professor![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-45.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-6.webp)](https://www.dataleadsfuture.com/i-used-autogen-graphflow-and-qwen3-coder-to-solve-math-problems-and-it-worked/) But right now, many commercial products that offer code sandboxes charge fees and limit resources. So in today’s tutorial, I’ll show you how to hook your agent up to a self-hosted Jupyter Server. You’ll get a powerful sandbox runtime with reusable context and solid computing power. This is a special, exclusive tutorial with full details—enough for you to master this core tech. So what are you waiting for? Let’s jump in. --- ## Environment Setup ### Build a Jupyter Kernel container The “code sandbox” your agent connects to relies on containerization for safety and environment isolation. So first, prepare a Docker image that runs Jupyter Server. The heart of any Docker container is the `Dockerfile`. To save you time, here’s the full content: ```text # Dockerfile.jupyter FROM python:3.13-slim-bookworm WORKDIR /app COPY requirements.txt /app/requirements.txt RUN pip install --no-cache-dir jupyter_kernel_gateway ipykernel numpy pandas sympy scipy --upgrade RUN pip install --no-cache-dir -r requirements.txt --upgrade EXPOSE 8888 ENV TOKEN="UNSET" CMD python -m jupyter kernelgateway \ --KernelGatewayApp.ip=0.0.0.0 \ --KernelGatewayApp.port=8888 \ --KernelGatewayApp.auth_token="${TOKEN}" \ --JupyterApp.answer_yes=true \ --JupyterWebsocketPersonality.list_kernels=true ``` This file uses `python:3.13-slim-bookworm` as the base image—not a pre-built Jupyter image—because we’ll customize the Jupyter environment later. I pulled essential dependencies out of `requirements.txt` and installed them separately. This maximizes Docker layer caching. Here’s the `requirements.txt` content: ```text matplotlib xlrd openpyxl ``` I included some basic Jupyter launch parameters. As we go, we’ll add more to build the complete Jupyter code sandbox. Once your `Dockerfile` is ready, run this command to build the image: ```shell docker build -t jupyter-server . ``` Don’t start the Jupyter container yet—we’ll explain why later. If your company can't use Docker Desktop due to licensing issues, don't worry—I've got an alternative solution for you. You can click here to read more: [A Quick Guide to Containerizing Agent Applications with PodmanAlternative solutions compatible with Docker SDK![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-48.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-8.webp)](https://www.dataleadsfuture.com/a-quick-guide-to-containerizing-agent-applications-with-podman/) ### Install the Autogen agent framework Most agent frameworks have moved Jupyter runtime support into paid offerings. Right now, Autogen is the only solid open option that supports Jupyter runtimes. To build agents, first install the `autogen-agentchat` package: ```shell pip install -U "autogen-agentchat" ``` To use containerized code executors, also install Autogen’s Docker client library: ```shell pip install "autogen-ext[docker-jupyter-executor]" ``` With the image built and Autogen installed, you’re ready to code. --- ## Using the Jupyter Code Sandbox ### Start with the recommended Docker API approach Let’s begin with the official API example to see how Autogen’s code executor works. Autogen has three key modules for Jupyter + Docker: `DockerJupyterCodeExecutor`, `DockerJupyterServer`, and `CodeExecutorAgent`. `DockerJupyterServer` uses the Docker API to start a container from a given image, mount directories, and store Jupyter connection info. `DockerJupyterCodeExecutor` handles all operations with the Jupyter Kernel API. Once it gets connection info from the server, you can submit and run code. `CodeExecutorAgent` is a special Autogen agent that pulls Python code from context and executes it. If you give it a `model_client`, it can even write its own code and reflect on results. ![The roles of different modules related to the Jupyter code sandbox.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-7.png) The roles of different modules related to the Jupyter code sandbox. Image by Author Now let’s build a code executor agent to test if this stateful Jupyter sandbox works. Remember the `jupyter-server` Docker image we built? Use it to initialize `DockerJupyterServer`. ```python server = DockerJupyterServer( custom_image_name="jupyter-server", expose_port=8888, token="UNSET", bind_dir="temp", ) ``` Then use that server to create a `DockerJupyterCodeExecutor` instance: ```python executor = DockerJupyterCodeExecutor( jupyter_server=server, timeout=600, output_dir=Path("temp") ) ``` Note: both `server` and `executor` mount your local `temp` folder into the container. Code can read/write files there, but inside the Jupyter kernel, the working directory is `/app`, not `temp`. Next, create the `CodeExecutorAgent`. Just pass the `executor` instance to the `code_executor` parameter. ```python code_executor = CodeExecutorAgent( "code_executor", code_executor=executor, ) ``` Now write a `main` method to test `coder_executor`. ```python async def main(): async with executor: code1 = TextMessage( content=dedent(""" ```python x = 1+2 print("Round one: The calculation for the value of x is done.") ``` """), source="user" ) response1 = await code_executor.on_messages(messages=[code1], cancellation_token=CancellationToken()) print(response1.chat_message.content) code2 = TextMessage( content=dedent(""" ```python print("Round two: Get the value of variable x again: x=", x) ``` """), source="user", ) response2 = await code_executor.on_messages(messages=[code2], cancellation_token=CancellationToken()) print(response2.chat_message.content) asyncio.run(main()) ``` To check stateful execution, call `code_executor` twice: First, define a variable `x` and compute something. Second, print `x`. In a command-line sandbox, this fails—the second run doesn’t know about `x`. But with Jupyter’s stateful kernel, the variable stays alive between calls (as long as you use the same executor): ![The code in the next round was able to access the variables from the previous round.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-8.png) The code in the next round was able to access the variables from the previous round. Image by Author I’ve already shown how this stateful sandbox helps agents solve hard problems. Read more here: [How I Crushed Advent of Code And Solved Hard Problems Using Autogen Jupyter Executor and Qwen3A detailed guide on using AI to generate code and solve puzzles automatically and quickly![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-42.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/jupyter_executor_cover.webp)](https://www.dataleadsfuture.com/how-i-crushed-advent-of-code-and-solved-hard-problems-using-autogen-jupyter-executor-and-qwen3/) This method—starting a Jupyter container from an image via code—is called “Docker out of Docker.” ### Problems with Docker out of Docker If you’re just testing Jupyter’s superpowers on your local machine, `DockerJupyterServer` works fine. But the big issue? The Jupyter Server actually starts on the same machine running your agent code. This breaks down if you need serious compute power, or plan to deploy to production: For data security or performance, companies often use powerful internal Jupyter Servers. If your data is gigabytes big, you need a server with tens of GBs of RAM—not your laptop. Things get worse if you containerize your agent app. Due to network isolation, your agent container might start a Jupyter container but fail to reach it. You wouldn’t run both agent and Jupyter on the same web server. Instead, deploy Jupyter on a dedicated compute server and let multiple agents share it—maximizing hardware use. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-9.png) You can let multiple agents access the same Jupyter Server. Image by Author For example, I rented a GPU server on vast.ai, set up JupyterLab, and want my agent to connect directly for data analysis. ### Let agents connect to the Jupyter Server directly By now it’s clear: to use external compute power, your agent must connect to a pre-deployed Jupyter Server—not spin up its own. You won’t find this solution anywhere online. So here’s today’s key reveal: how to connect your multi-agent app to a self-hosted Jupyter Kernel server—for low cost (vs. Azure/Claude) and high compute power. Go back to the section where we launched Jupyter from a Docker image. Remember: `DockerJupyterServer` saves connection info after startup, and `DockerJupyterExecutor` uses that to connect. What if we skip `DockerJupyterServer` and make `DockerJupyterExecutor` connect directly to a standalone Jupyter Server? Check the `DockerJupyterExecutor` source code: ```python class DockerJupyterCodeExecutor(CodeExecutor, Component[DockerJupyterCodeExecutorConfig]): ... def __init__( self, jupyter_server: Union[JupyterConnectable, JupyterConnectionInfo], kernel_name: str = "python3", timeout: int = 60, output_dir: Path | None = None, ): ... if isinstance(jupyter_server, JupyterConnectable): self._connection_info = jupyter_server.connection_info elif isinstance(jupyter_server, JupyterConnectionInfo): self._connection_info = jupyter_server ``` At init, it sets a `_connection_info` member. - If you pass a `JupyterConnectionInfo` instance, it uses that directly. - If you pass a `DockerJupyterServer` instance, it reads `.connection_info` from it. Earlier, we passed a `DockerJupyterServer` instance. Now let’s try passing `JupyterConnectionInfo` instead. First, find your Jupyter connection details. If you launched from our image, it’s: `host='127.0.0.1'`, `use_https=False`, `port=8888`, `token='UNSET'`. For internal or vast.ai servers, check your browser’s address bar. ![You can get the Jupyter connection info from the browser’s address bar.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-10.png) You can get the Jupyter connection info from the browser’s address bar. Image by Author Now update the `DockerJupyterCodeExecutor` init, pass `JupyterConnectionInfo` directly: ```python executor = DockerJupyterCodeExecutor( jupyter_server=JupyterConnectionInfo( host='127.0.0.1', use_https=False, port=8888, token='UNSET' ), timeout=600, output_dir=Path("temp"), ) ``` When we re-run `main`, it crashes—because I’m trying to connect to a container that isn’t running yet. ### Manage instances gracefully with Docker Compose To test our updated agent, we must first start the Jupyter container. If you know Docker well, just run `docker run`. ```shell docker run -d -p 8888:8888 --volume temp:/app --name jupyter-server jupyter-server ``` Again, I recommend [**DataCamp’s ‘Introduction to Docker’ course**](https://datacamp.pxf.io/jekVAv?ref=dataleadsfuture.com) to master the basics fast. I’ll level up: when starting, I’ll mount the local `temp` folder into the container’s `/app` workdir—so code can read/write files. That command feels messy, right? Honestly, I haven’t used it in ages. I use Docker Compose instead. Docker Compose manages groups of related containers. For single-image setups, it’s super clean: write a `docker-compose.yml` file in your project folder. ```text version: "3.8" services: jupyter: image: jupyter-server container_name: jupyter-server ports: - "8888:8888" volumes: - ./temp:/app networks: - docker_executor networks: docker_executor: driver: bridge ``` Run `docker compose up -d` to start, and `docker compose down` to stop. ![Use Docker Compose to manage your Jupyter container.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-11.png) Use Docker Compose to manage your Jupyter container. Image by Author After starting Jupyter Server, re-run `main`. To test stateful execution, put a simple CSV in `temp` and read it: ```python async def main2(): async with executor: code1 = TextMessage( content=dedent(""" ```python from pathlib import Path import pandas as pd file_path = Path("superstore.csv") df = pd.read_csv(file_path) print(df.iloc[:5, :6].head()) ``` """), source="user", ) response1 = await code_executor.on_messages(messages=[code1], cancellation_token=CancellationToken()) print(response1.chat_message.content) code2 = TextMessage( content=dedent(""" ```python region_sales_sum = df.groupby("Region", as_index=False)["Sales"].sum() print(region_sales_sum) ``` """), source="user", ) response2 = await code_executor.on_messages(messages=[code2], cancellation_token=CancellationToken()) print(response2.chat_message.content) asyncio.run(main2()) ``` In this `main`, I first load and preview the CSV. Then in a second code block, I group by a column and sum values. ![In the first code snippet, read the CSV data, and in the second snippet, do the calculations.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-12.png) In the first code snippet, read the CSV data, and in the second snippet, do the calculations. Image by Author See? The file loads fine in Jupyter Server, and the code runs incrementally in the stateful environment. ### Tune the Jupyter image to reclaim idle resources When using Docker API, container resources are auto-cleaned on exit (thanks to `async with`). But with standalone deployment, every new `DockerJupyterCodeExecutor` connection spawns a new Jupyter Kernel. Even after disconnecting, the kernel stays—wasting memory. So we must tweak the Jupyter image’s `Dockerfile` to auto-cleanup idle kernels. Add these flags to the `Jupyter KernelGateway` launch command: ```text CMD python -m jupyter kernelgateway \ --KernelGatewayApp.ip=0.0.0.0 \ --KernelGatewayApp.port=8888 \ --KernelGatewayApp.auth_token="${TOKEN}" \ --JupyterApp.answer_yes=true \ --JupyterWebsocketPersonality.list_kernels=true \ --MappingKernelManager.cull_idle_timeout=1800 \ --MappingKernelManager.cull_interval=300 \ --MappingKernelManager.cull_connected=False \ --MappingKernelManager.cull_busy=False \ ``` Key settings: - `cull_idle_timeout`: kill kernel after X seconds idle - `cull_interval`: check for idle kernels every X seconds - `cull_connected`: reclaim kernels with active connections? - `cull_busy`: force-kill busy kernels? With this, disconnected clients free up resources automatically. No more OOM crashes from long-running servers. Remember to rebuild the image so `Dockerfile` changes take effect. --- ## A Simple Multi-Agent Project Demo By now, you’ve mastered the Jupyter code sandbox setup. We even tested stateful execution with hand-written code blocks. But real projects need LLMs—not humans—to generate Python code step-by-step based on user tasks. So let’s expand: build a system where the LLM breaks down user requests into incremental Python steps. Besides `code_executor`, add two new agents: - `task_planner` splits complex user questions into subtasks. It outputs one new step at a time. - `code_writer` turns each subtask into executable Python code and sends it to `code_executor`. ![Use an iterative way to break a big problem into small code snippets and run them one by one.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-13.png) Use an iterative way to break a big problem into small code snippets and run them one by one. Image by Author Here’s `task_planner`’s code: ```python SYSTEM_PROMPT = dedent(""" You are the task planning helper in the team, good at breaking down complex user requests into smaller sub-tasks that can be done with Python code. ## Duties 1. **Only split tasks**, don’t write code or do the sub-tasks yourself. 2. **Make just one sub-task at a time**, don’t skip steps or merge different steps together. 3. **Think about the context**, use the results from earlier steps to make new and reasonable sub-tasks. 4. **Create tasks step by step**, keep breaking things down until the user’s original request is fully answered. 5. When all sub-tasks are done, **make a summary report based on the work history**. 6. At the very end, output "**TERMINATION**" as the finish signal. """) planner = AssistantAgent( "task_planner", model_client=model_client, system_message=SYSTEM_PROMPT, ) ``` Here’s `code_writer`’s code: ```python SYSTEM_PROMPT = dedent(""" You’re a code helper in the team, good at writing Python code that can run in a stateful Jupyter Kernel based on the task you need to do. ## Responsibilities 1. **Understand the task**: Clearly understand the analysis or data processing request you’re given. 2. **Write code step by step**: Build the code in small, growing steps, making full use of the Jupyter Kernel’s stateful feature (meaning variables, data, and state stay between code blocks), and avoid running the same thing more than once. 3. **Show the output clearly**: Make sure each piece of code shows or returns its result clearly so the team can see and check it. 4. **Follow code format rules**: All Python code must be wrapped in Markdown code blocks like ` ```python ` to keep it easy to read and run. 5. **Reuse context**: Let later code blocks use variables, data frames, models, and other things you set up earlier, without loading or starting them again. ## Examples When you write Python code, wrap it in a markdown python code block: ```python x = 3 ``` You can reuse the variable in another code block: ```python print(x) ``` """) code_writer = AssistantAgent( "code_writer", model_client=model_client, system_message=SYSTEM_PROMPT, ) ``` Since we solve problems iteratively, we use Autogen’s `RoundRobinGroupChat` to loop until the user’s question is answered: ```python team = RoundRobinGroupChat( [planner, code_writer, code_executor], termination_condition=combine_term ) ``` Test it with a `main` method using Kaggle’s `superstore` dataset: ```python if __name__ == "__main__": async def main(): async with executor: await Console( team.run_stream(task="Read the superstore.csv file and find the total sales for each region.") ) asyncio.run(main()) ``` See? The agent runs code step-by-step, gets the final result, and even adds insights. Jupyter code sandboxes truly unlock agent potential: ![The agent solved the needed metrics step by step and shared its own insights.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-14.png) The agent solved the needed metrics step by step and shared its own insights. Image by Author --- ## Can LangChain or Other Frameworks Use Jupyter Code Sandboxes? So far, we’ve used Autogen to harness Jupyter sandboxes—and tests prove their power for complex tasks. But maybe you use LangChain, CrewAI, or another framework. Can they connect to Jupyter sandboxes as easily as Autogen? Yes! ### Create an executor function At its core, we use `DockerJupyterCodeExecutor` to talk to Jupyter Server. `CodeExecutorAgent` isn’t required: you can wrap the executor in a simple function and expose it as a tool. Take LangChain as an example. Code goes in `langchain_with_jupyter_executor.py`. Initialize `executor` as before, using `JupyterConnectionInfo` to connect to your standalone server. Then create an `execute_code` function and mark it as a LangChain tool with `@tool`: ```python @tool async def execute_code(code: str) -> str: """ Use the Jupyter code executor to run your Python code. The runtime environment keeps its state, so you can run code step by step. reuse variables from earlier code blocks, and avoid writing the same code again. :param code: Code waiting to be run, only the code itself, no Markdown syntax :return: The result of the code execution. """ code_blocks = [CodeBlock(code=code, language="python")] code_result = await executor.execute_code_blocks(code_blocks, cancellation_token=CancellationToken()) return code_result.output ``` Important: LLMs often output code wrapped in Markdown code blocks. But `executor` needs raw Python. Mention this clearly in the function docstring. ### Make LangChain use this tool Now create a LangChain `model client` and agent. In the `system_prompt`, tell it: “You can write Python code and send it to the `execute_code` tool.” ```python model = ChatOpenAI( model="qwen3-next-80b-a3b-instruct", api_key=os.getenv("OPENAI_API_KEY"), base_url=os.getenv("OPENAI_BASE_URL"), temperature=0.1, top_p=0.85, ) agent = create_agent( model=model, tools=[execute_code], system_prompt=dedent(""" You are a data analysis assistant, good at solving user questions with Python code. You use the `execute_code` tool to run the code and summarize the results as the answer. """) ) ``` Test it with a simple `main`: ```python async def main(): async with executor: result = await agent.ainvoke( {"messages": [ {"role": "user", "content": "Calculate the value of the 14th Fibonacci number."} ]} ) for msg in result['messages']: print(msg.content) asyncio.run(main()) ``` Success! The agent wrote Python code based on the user’s request and got the answer. ```text Calculate the value of the 14th Fibonacci number. 377 The 14th Fibonacci number is 377. ``` You can do the same with LangGraph or any agent framework: wrap executor calls in a tool function, then use function calling to trigger it. ![LangChain can run code by calling the executor with a function call.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-15.png) LangChain can run code by calling the executor with a function call. Image by Author Note: `DockerJupyterCodeExecutor` itself uses Jupyter KernelGateway’s REST API. If you know Jupyter APIs well, you could build a custom `CodeExecutor` for any framework. --- ## Conclusion Past tutorials proved that running agent-generated code in a stateful Jupyter sandbox unlocks huge potential for solving complex user problems. But most multi-agent frameworks either lack this feature or only connect to cloud-based, resource-limited, commercial sandboxes. Today, I showed you exactly how to connect Autogen to a self-hosted Jupyter Server code sandbox. I broke it down from multiple deployment angles so you fully master this technique. And you’re not limited to Autogen. Any agent framework can gain this power through function calling. Jupyter sandbox connectivity works across frameworks. Remember: incremental, iterative code execution in a Jupyter sandbox is a foundational skill for **building deep data analysis agents**. In future posts, I’ll cover other core technologies. Together, we’ll build a stable, enterprise-grade deep data analysis agent system. Stay tuned! --- Here’s today’s project source code: [agentic-ai-playground/09\_Decrypt\_Jupyter\_Code\_Executor at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-21.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-19)](https://github.com/qtalen/agentic-ai-playground/tree/main/09%5FDecrypt%5FJupyter%5FCode%5FExecutor?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### How To Use DeepSeek-OCR And Docling For PDF Parsing URL: https://www.dataleadsfuture.com/how-to-use-deepseek-ocr-and-docling-for-pdf-parsing/ Last updated: 2026-08-06T01:53:42.000Z This is a hands-on guide that explores how to use DeepSeek-OCR with docling for PDF parsing inside an agent application. By reading this, you will learn how to use the DeepSeek-OCR model in real code. I will also compare DeepSeek-OCR's results with a traditional OCR model so you can see how good it actually is. You will find the full source code for this project at the end. --- ## Introduction DeepSeek-OCR has been making waves lately — articles and videos are everywhere. The idea of using Contexts Optical Compression instead of text tokens sounds amazing. But once you read those posts carefully, you find they're basically reposts of DeepSeek's official blog charts with a lot of praise like “optical compression will change how LLMs understand context” and other blah blah. They rarely explain how to use this model in a real project. So, is this one of those “nice in the lab but useless in the real world” things? Not for us. Today I’ll walk you through how to use DeepSeek-OCR inside docling’s `VlmPipeline` to parse PDFs. For our test file, I picked NVIDIA FY2026 Q2 Financial Report. Here’s a screenshot of the parsed result: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image.png) Compare the original PDF text (top) with the parsed Markdown content (bottom). Image by Author I’ll also add [cognee](https://docs.cognee.ai/getting-started/introduction?ref=dataleadsfuture.com) to run Agentic RAG on the parsed text. Even though DeepSeek-OCR focuses on the optical compression concept, I believe accurate OCR for PDFs is the basic requirement for any VLM. The DeepSeek team said in their official blog that they used PaddleOCR for data labeling. So later in this post, I’ll compare PDF parsing results from DeepSeek-OCR and PaddleOCR so you can see for yourself if DeepSeek-OCR’s performance lives up to the hype. Buckle up, let’s go. --- ## Environment Setup ### Install dependencies Before coding, let’s install the dependencies we need for docling and DeepSeek-OCR. I’m using `vllm server` for deploying the DeepSeek-OCR model, so I don’t need vllm-related packages locally. You only need to install `docling`, `docling[vlm]`, and `hf-xet`. For Agentic RAG with cognee, install `cognee>=0.3.6` and `starlette>=0.48`. Watch the version number — starlette versions below `0.48` will make cognee throw errors. We also need PaddleOCR for comparison. To make it easier, I’m using the `onnxruntime` version of RapidOCR. So add those two dependencies too. Don’t worry — my project’s `pyproject.toml` already lists them all. Just run: ```shell pip install --upgrade -e . ``` and you’re set. ### Configure environment variables I’m using DeepSeek-OCR deployed in OpenAI API compatible mode (though you can get a provider that hosts the model too). So in the `.env` file, I set up: ```text OCR_MODEL="deepseek-ai/DeepSeek-OCR" OCR_API_KEY= OCR_BASE_URL= ``` I use `OCR_API_KEY` and `OCR_BASE_URL` to distinguish from regular LLM clients, but you can change those in code. To use cognee, you also need LLM and embedding model configs. This isn’t the focus here, so check their docs for details. ```text # Cognee LLM Provider setup LLM_PROVIDER="openai" LLM_MODEL="openai/Qwen/Qwen3-30B-A3B-Instruct-2507" LLM_API_KEY= LLM_ENDPOINT= LLM_RATE_LIMIT_ENABLED="true" LLM_RATE_LIMIT_REQUESTS="600" LLM_RATE_LIMIT_INTERVAL="60" # Cognee Embedding model setup EMBEDDING_PROVIDER="custom" EMBEDDING_MODEL="openai/BAAI/bge-m3" EMBEDDING_DIMENSIONS="1024" EMBEDDING_API_KEY= EMBEDDING_ENDPOINT= EMBEDDING_RATE_LIMIT_ENABLED="true" EMBEDDING_RATE_LIMIT_REQUESTS="2000" EMBEDDING_RATE_LIMIT_INTERVAL="60" ``` --- ## Integrating DeepSeek-OCR And Docling ### Module design Sure, you can find some PDF parsing examples with DeepSeek-OCR online, but most are just notebook experiments. I want our parser to be proper and ready for enterprise use. So I’ll write it with a clean, modular design. For enterprise, rewriting everything from scratch — splitting PDFs, converting them to images, passing to a vision model — is a waste. We should use what we already have. In this case, that’s docling. Here’s the flow chart for our module: ![How the OCR module works.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/DeepSeek-OCR_with_docling-OCR.drawio.png) How the OCR module works. Image by Author Cognee is optional — swap in any RAG system you like. ### Start coding We’ll organize code with OOP, inside `ocr_agentic_rag.py`. First, define an `OCRAgenticRAG` class. All parsing happens here. ```python import os import asyncio from pathlib import Path from tempfile import gettempdir from dotenv import load_dotenv from docling.datamodel.base_models import InputFormat from docling.datamodel.pipeline_options import ( VlmPipelineOptions ) from docling.datamodel.pipeline_options_vlm_model import ( ApiVlmOptions, ResponseFormat ) from docling.document_converter import ( DocumentConverter, PdfFormatOption ) from docling.pipeline.vlm_pipeline import VlmPipeline import cognee from cognee.infrastructure.databases.vector.embeddings.config import EmbeddingConfig from common.utils.project_path import get_project_root, get_current_directory load_dotenv(get_project_root() / ".env") class OCRAgenticRAG: ... ``` The two key methods are `_openai_compatible_vlm_options` and `_get_docling_converter`. Since we let docling’s pipeline talk to DeepSeek instead of using `requests`, `_openai_compatible_vlm_options` sets up the VLM client: ```python class OCRAgenticRAG: ... @staticmethod def _openai_compatible_vlm_options( model: str = "", prompt: str = "Convert these pdf pages to markdown.", response_format: ResponseFormat = ResponseFormat.MARKDOWN, base_url: str = "", temperature: float = 0.7, max_tokens: int = 4096, api_key: str = "", skip_special_token = False, ): ocr_model = model or os.getenv("OCR_MODEL") headers = {} if api_key: headers["Authorization"] = f"Bearer {api_key}" headers["Content-Type"] = "application/json" options = ApiVlmOptions( url=f"{base_url}/chat/completions", params=dict( model=ocr_model, max_tokens=max_tokens, skip_special_token=skip_special_token, ), headers=headers, prompt=prompt, timeout=90, scale=1.0, temperature=temperature, response_format=response_format, ) return options ``` Next, `_get_docling_converter` configures `pipeline_options` so docling uses that VLM setup for PDF processing. ```python class OCRAgenticRAG: ... def _get_docling_converter( self, api_key: str = "", base_url: str = "", ) -> DocumentConverter: pipeline_options = VlmPipelineOptions( enable_remote_services=True ) pipeline_options.vlm_options = self._openai_compatible_vlm_options( api_key=api_key, base_url=base_url ) doc_converter = DocumentConverter( format_options={ InputFormat.PDF: PdfFormatOption( pipeline_options=pipeline_options, pipeline_cls=VlmPipeline, ) } ) return doc_converter ``` Once we set up the docling PDF converter, `_ocr_pdf` parses PDFs into Markdown `md` files. It supports multiple files at once. To track results, I save those Markdown files in a temp folder. ```python class OCRAgenticRAG: ... def _ocr_pdf(self, source_data: str | list[str]) -> list[str]: if not isinstance(source_data, list): source_data = [source_data] output_files = [] for source_file in source_data: result = self.converter.convert(source_file) markdown_str = result.document.export_to_markdown() source_filename = result.input.file.stem out_file = self._write_to_file(self.temp_dir, source_filename, markdown_str) output_files.append(out_file) return output_files ``` Finally, add the entry method for handling files. ```python class OCRAgenticRAG: ... async def add(self, files: str | list[str]): temp_files = self._ocr_pdf(files) print("All the PDF files have been successfully parsed.") ``` Now docling and DeepSeek-OCR can work together. If you also want Agentic RAG with cognee, add this code: ```python class OCRAgenticRAG: ... @staticmethod async def clear(): await cognee.prune.prune_data() await cognee.prune.prune_system(metadata=True) async def add(self, files: str | list[str]): temp_files = self._ocr_pdf(files) print("All the PDF files have been successfully parsed.") await self.clear() await cognee.add(temp_files) await cognee.cognify() @staticmethod async def search(query: str) -> str: results = await cognee.search( query_text=query ) return "\n".join([str(result) for result in results]) ``` ### Test the code Next, let’s write a simple `main` method to check how this parsing module works. ```python if __name__ == "__main__": ocr_rag = OCRAgenticRAG(temp_dir=get_current_directory()/"temp") async def main(): source_dir = get_current_directory() / "temp" pdf_files = [ source_dir / "NVIDIAAn.pdf" ] await ocr_rag.add(pdf_files) result = await ocr_rag.search(query="How much was Nvidia’s revenue in Q2 Fiscal 2025?") print(result) asyncio.run(main()) ``` We’ll parse the NVIDIA FY2026 Q2 Financial Report since reading financial reports is a classic AI agent use case. We’ll search with a question to see if cognee indexed the right text: ![I asked for the data for the 2025 fiscal year, but it gave me the results for the 2026 fiscal year.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-1.png) I asked for the data for the 2025 fiscal year, but it gave me the results for the 2026 fiscal year. Image by Author Wait — the PDF definitely has NVIDIA’s FY2025 Q2 revenue info. Why can’t cognee find it? --- ## Evaluating DeepSeek-OCR’s PDF parsing Since RAG missed something, I wanted to check the quality of DeepSeek-OCR’s Markdown output. Comparing with the original PDF, I found many mistakes in tables. Multi-header tables had missing rows and columns. ![DeepSeek-OCR made a lot of mistakes and lost data when handling multi-headers and across-page parts in file parsing.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-2.png) DeepSeek-OCR made a lot of mistakes and lost data when handling multi-headers and across-page parts in file parsing. Image by Author But hold on — the docs say DeepSeek-OCR reaches 97% accuracy. It shouldn’t be this bad. Maybe this PDF is just hard to parse? I decided to make a control group. Since the DeepSeek team used PaddleOCR for labeling, I’ll parse the file with PaddleOCR to see if the document itself is tricky. Docling doesn’t support PaddleOCR directly, but RapidOCR in docling is basically PaddleOCR with `onnxruntime`. Accuracy is a bit lower, but it’s fine for testing. The official docs already show how. My code is in `paddle_ocr_docling.py`. ```python from docling.datamodel.base_models import InputFormat from docling.datamodel.pipeline_options import ( PdfPipelineOptions, RapidOcrOptions, ) from docling.document_converter import DocumentConverter, PdfFormatOption from common.utils.project_path import get_current_directory def main(): pdf_file = get_current_directory() / "temp/NVIDIAAn.pdf" pipeline_options = PdfPipelineOptions() pipeline_options.do_ocr = True pipeline_options.do_table_structure = True pipeline_options.table_structure_options.do_cell_matching = True ocr_options = RapidOcrOptions( force_full_page_ocr=False, ) pipeline_options.ocr_options = ocr_options converter = DocumentConverter( format_options={ InputFormat.PDF: PdfFormatOption( pipeline_options=pipeline_options, ) } ) doc = converter.convert(pdf_file).document md = doc.export_to_markdown() with open(get_current_directory() / "temp/NVIDIAAn_rapid.md", "w", encoding="utf8") as f: f.write(md) if __name__ == "__main__": main() ``` Let’s check the Markdown from RapidOCR, especially the tables. ![RapidOCR’s table data parsing is super accurate.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/11/image-3.png) RapidOCR’s table data parsing is super accurate. Image by Author Looks like RapidOCR’s tables are much more accurate, with no major missing or misplaced data. --- ## Conclusion DeepSeek’s popularity means that whenever the team drops something new, the internet goes wild. The Contexts Optical Compression idea in DeepSeek-OCR, if proven viable, would spark another wave of innovation for LLMs. But most online content is just theory talk. Few actually show you how to use DeepSeek-OCR in a real project. This tutorial aimed to show how to integrate DeepSeek-OCR with docling and other open-source tools to parse PDFs. Still, as an OCR model, DeepSeek-OCR’s table handling needs work for both text-based and scanned PDFs. This limits its use in data science. I hope to see more practical experiments on where DeepSeek-OCR shines. If you have thoughts, drop me a comment. --- ## Project Source Code Here’s the source code of this tutorial — feel free to read or use it: [Source code repository on GitHub.](https://github.com/qtalen/agentic-ai-playground/tree/main/08%5FDeepSeek%5FOCR%5FAgentic%5FRAG?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### How I Crushed Advent of Code And Solved Hard Problems Using Autogen Jupyter Executor and Qwen3 URL: https://www.dataleadsfuture.com/how-i-crushed-advent-of-code-and-solved-hard-problems-using-autogen-jupyter-executor-and-qwen3/ Last updated: 2026-04-22T08:23:14.000Z In today’s tutorial, I will use Autogen’s `docker-jupyter-executor` runtime with Qwen3’s newest `max` model to try finishing the Advent of Code challenge quickly. I aim to demonstrate that combining LLM code generation with a stateful Python runtime can effectively solve extremely complex algorithmic problems. As usual, I will share the full project source code for you to check. You can find it in the reference section. --- ## Course Background You have probably heard of [Advent of Code (AOC)](https://adventofcode.com/?ref=dataleadsfuture.com). It is a fun programming challenge that claims to help beginners practice a programming language. The puzzles are really hard. Every year, I struggle and only finish the first few days. I was not happy about that. This year, there is still one month before the Advent of Code starts, but I have done all the prep work. New monitor, new IDE, new keyboard, and a new agent tool. Yes, I do not plan to use my brain to solve the problems this year. Like AlphaGo, I want to build an agent. I will let AI read the puzzle, write the code, and get the result all by itself. My job will be making coffee and sitting at my desk waiting. It worked. I tested with past challenges and started getting stars faster than I had time to read the problems. My cost was only some tokens. ![My agent scored 47 stars in the 2024 Advent of Code challenge.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image.png) My agent scored 47 stars in the 2024 Advent of Code challenge. Image by Author I even allow users to enter Part Two of an AoC problem through multi-turn conversation, so the agent can keep solving. ![The agent supports continuous problem solving through multi turn conversation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/aoc_2023_day_8.gif) The agent supports continuous problem-solving through multi-turn conversation. Image by Author And it is not just for Advent of Code. This agent can also run data analysis or other tasks you can imagine. ![I used this agent for abnormal transaction detection.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/fraud_detection.gif) I used this agent for abnormal transaction detection. Image by Author ### How did this all happen In today’s tutorial, you will see: - I will follow the ReAct pattern to build a single-agent app that solves complex challenges by planning sub-steps one at a time. - Each sub-step depends on the main task and previous results, so the LLM can adjust mistakes anytime. - Each sub-step uses Python code to solve the puzzle and uses Jupyter as the runtime to get intermediate results. - The agent relies on the stateful Jupyter kernel, so it can reflect on previous results and adjust the next steps until it finds the final answer. The effect is amazing. ### Why this works well In my last post, we tried building a multi-agent system for math problems. You can read it here: [I Used Autogen GraphFlow and Qwen3 Coder to Solve Math Problems — And It WorkedMore reliable than your math professor![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-40.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-2.webp)](https://www.dataleadsfuture.com/i-used-autogen-graphflow-and-qwen3-coder-to-solve-math-problems-and-it-worked/) That system worked well, but not perfectly. It worked by letting a reasoning agent plan all steps at once and then sending them to a coding agent to write Python code. This caused problems. For exploratory tasks like reading a file and then deciding what to do based on its content, the system could not handle it. If the code failed during execution, the whole Python file had to be regenerated to find the error and adjust it. This was not flexible. Think about how humans handle challenging tasks like data analysis or ML modeling. We write some code, run it, see if the result matches expectations, then decide what to write next. That is why Jupyter is so popular in data science. So why not use Jupyter as the Python runtime? Of course, we can. That is what we will do today. We will generate a small bit of code each time we run it, then move forward until we reach the goal. --- ## Preparation ### Make sure the LLM doesn't include AOC 2024 data First, we need to ensure that the LLM we use hasn't been trained on Advent of Code 2024 data. This is the foundation of this project. ![Qwen3 max model does not include any information about AOC 2024.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image-7.png) Qwen3 max model does not include any information about AOC 2024\. Image by Author Also, I don't plan to give the agent any way to get information from the internet. ### Build a Jupyter container Since we will use Jupyter as the runtime, we need to set it up before the course starts. I will use a Docker container to isolate Jupyter so that bad LLM code will not break the system. `Dockerfile` looks like this: ```text FROM python:3.13-slim-bookworm WORKDIR /app COPY requirements.txt /app/requirements.txt RUN pip config set global.index-url https://mirrors.aliyun.com/pypi/simple/ && \ pip install --no-cache-dir jupyter_kernel_gateway ipykernel numpy pandas sympy scipy --upgrade RUN pip install --no-cache-dir -r requirements.txt --upgrade EXPOSE 8888 ENV TOKEN="UNSET" CMD python -m jupyter kernelgateway \ --KernelGatewayApp.ip=0.0.0.0 \ --KernelGatewayApp.port=8888 \ --KernelGatewayApp.auth_token="${TOKEN}" \ --JupyterApp.answer_yes=true ``` `requirements.txt` looks like this: ```text matplotlib xlrd openpyxl pdfplumber reportlab ``` I install rarely changed dependencies and often changed dependencies separately to use Docker layer cache for faster builds. Autogen uses Docker SDK to control the start and stop of the container, so I did not set up Jupyter auth. This makes the runtime call easier, but it is not safe for production. Then we build the image and name it `jupyter-server` for later. ```shell docker build -t jupyter-server . ``` ### Test connectivity with Autogen After building the image, we need to test with Autogen to see if running code in Jupyter works. We must install `autogen-ext[docker-jupyter-executor]` and `nbclient`. Do not worry. I already added these to `pyproject.toml` So you just run `pip install --upgrade -e .`. Before starting, we need to initialize a `DockerJupyterServer` module. This uses Docker SDK to start a container from the Jupyter image. We will use this today. ```python jupyter_server = DockerJupyterServer( custom_image_name="jupyter-server:latest", expose_port=8888 ) ``` This way of using the Docker SDK to manage Jupyter images is called the Docker out-of-Docker method. But this method has some problems when used in production or when connecting to a separately deployed Jupyter service. Don’t worry—my latest tutorial has already solved these issues for you: [Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent SkillsUse Jupyter code executor to help your agent finish tasks in a smarter way![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-44.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-5.webp)](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) There are three ways to use Jupyter runtime. First, extract the Python code generated by the LLM, run it manually through the Jupyter executor, and get the result. ```python async def main_1() -> None: async with jupyter_server: async with DockerJupyterCodeExecutor(jupyter_server) as executor: code_blocks = [CodeBlock(code="print('hello world!')", language="python")] code_result = await executor.execute_code_blocks(code_blocks, cancellation_token=CancellationToken()) print(code_result) ``` Note that `DockerJupyterCodeExecutor` is stateful, so in an `async with` scope repeated calls reuse previous variables without regenerating them. Second use `PythonCodeExecutionTool` to execute code and return results. ```python async def main_2() -> None: async with jupyter_server: async with DockerJupyterCodeExecutor(jupyter_server) as executor: tool = PythonCodeExecutionTool(executor) agent = AssistantAgent("assistant", model_client=model_client, tools=[tool]) result = await agent.run(task="What is the 10th Fibonacci number? Use Python to calculate it.") print(result.messages[-1].content) ``` This uses the agent’s function call ability. If your agent needs to do many jobs and code execution is just one part, use this. Third use `CodeExecutorAgent` to execute code. ```python async def main_3() -> None: async with jupyter_server: async with DockerJupyterCodeExecutor(jupyter_server) as executor: code_executor_agent = CodeExecutorAgent("code_executor", code_executor=executor) task = TextMessage( content=""" ```python a = 3 ``` """, source="user" ) response = await code_executor_agent.on_messages([task], CancellationToken()) print(response.chat_message) task_2 = TextMessage( content=""" ```python print(a) ``` """, source="user" ) response_2 = await code_executor_agent.on_messages([task_2], CancellationToken()) print(response_2.chat_message) ``` In a multi-agent system, if you want a dedicated agent for code execution and reflection, this is good. For example, in my last tutorial I used `CodeExecutorAgent` in an `Autogen GraphFlow` to handle code execution. ![Multi-agent system for generating code and getting results built with Autogen GraphFlow. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image-3.png) Multi-agent system for generating code and getting results built with Autogen GraphFlow. Image by Author --- ## Let’s Start With the Jupyter runtime ready, we can look at today’s project. ### Architecture design Advent of Code is hard. No LLM can plan the whole logic up front. So we will plan one step, run the code, see the result, then plan the next. So the loop becomes think, act, observe, think again. Sounds familiar. Yes, this is the famous ReAct agent design. ![Diagram of a typical ReAct agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image-4.png) Diagram of a typical ReAct agent. Image by Author Since ReAct only needs one agent, we will build a single-agent app. The agent will use the user request and the previous result to plan the current step, then write a Python snippet to get the intermediate result. With a single agent app, it fits to use `PythonCodeExecutorTool` for running code. Unlike traditional generate and run code, here we plan one step and get only an intermediate result. In this case direct Python runtime does not work well. The best way is to send code to a Jupyter kernel, which saves variables and results. Our single-agent app architecture looks like this: ![Architecture diagram of our single agent app.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image-5.png) Architecture diagram of our single agent app. Image by Author ### Write agent code With goals and design set, it is coding time. Using Docker means we need to manage context and container lifecycle. I do not want the caller to start or stop Docker each time. Code execution is the agent’s duty, not the caller’s. I also want to keep the `Autogen AssistantAgent` API so the agent stays general. So I will wrap its init and call it as a new Agent. The agent and Jupyter runtime must allow generated code to read files. So I will mount a folder in the Docker container and put user-uploaded files in it. ```python class AOCAssistant: ... @staticmethod def _copy_file( file_name: str | None = None, file_path: Path | str | None = None, ) -> Path | str | None: if file_path is None: return None if file_name is None: file_name = Path(file_path).name dst_path = BINDING_DIR / file_name shutil.copy2(file_path, dst_path) return file_name ``` The Agent will manage `DockerJupyterServer` and `DockerJupyterCodeExecutor` lifecycle. ```python class AOCAssistant: ... async def start(self): await self._executor.start() async def stop(self): await self._model_client.close() await self._executor.stop() await self._jupyter_server.stop() async def __aenter__(self) -> "AOCAssistant": await self.start() return self async def __aexit__(self, exc_type, exc_val, exc_tb): await self.stop() def _init_jupyter_docker(self) -> None: self._jupyter_server = DockerJupyterServer( custom_image_name="jupyter-server:latest", expose_port=8888, bind_dir=BINDING_DIR, ) self._executor = DockerJupyterCodeExecutor( jupyter_server=self._jupyter_server, timeout=600) ``` I implemented `__aenter__` and `__aexit__` , so you can manage resources with `async with`. Next, init LLM client and `AssistantAgent` , bind the `CodeExecutor` as a tool to the Agent. ```python class AOCAssistant: ... def _init_assistant(self) -> None: self._model_client = OpenAILikeChatCompletionClient( model=self._model_name, temperature=0.5, top_p=0.85, ) tool = PythonCodeExecutionTool(self._executor) self._agent = AssistantAgent( 'assistant', model_client=self._model_client, tools=[tool], model_client_stream=True, system_message=SYS_PROMPT, max_tool_iterations=30, ) ``` I used the newest `Qwen3-max` model. Open source `qwen3-next-80b-a3b-instruct` is also good. I set `temperature` to 0.5 for some creativity in final results and `top_p` to 0.85 for serious planning and coding. I need ReAct style iteration, so I set `max_tool_iterations` in `AssistantAgent`. In Autogen, this lets the agent iterate based on `tool_calls`. It stops when it hits the max. Finally, to keep our custom Agent API the same as `Autogen AssistantAgent` I implemented `run` and `run_stream`. ```python class AOCAssistant: ... async def run( self, *, task: str | BaseChatMessage | Sequence[BaseChatMessage] | None = None, cancellation_token: CancellationToken | None = None, file_name: str | None = None, file_path: Path | str | None = None, ) -> TaskResult: async for message in self.run_stream( task=task, cancellation_token=cancellation_token, file_name=file_name, file_path=file_path, ): if isinstance(message, TaskResult): return message raise ValueError("No task result output.") async def run_stream( self, *, task: str | BaseChatMessage | Sequence[BaseChatMessage] | None = None, cancellation_token: CancellationToken | None = None, file_name: str | None = None, file_path: Path | str | None = None, ) -> AsyncGenerator[BaseAgentEvent | BaseChatMessage | TaskResult, None]: file_name = self._copy_file(file_name, file_path) input_messages = [] if isinstance(task, str): input_messages.append(TextMessage( source="user", content=task )) elif isinstance(task, BaseChatMessage): input_messages.append(task) if file_name is not None: input_messages.append(TextMessage( source="user", content=f"The input file is `{file_name}`" )) async for message in self._agent.run_stream( task=input_messages, cancellation_token=cancellation_token): yield message ``` `run` just calls `run_stream` and returns `TaskResult`. `run_stream` copies user files to the mounted directory, rebuilds `input_messages` adds file info, then calls `AssistantAgent.run_stream` to get LLM streaming output. ### Write the prompt This project needs the agent to plan sub-tasks step by step, write correct Python code, iterate based on results, and give a good final output. So the prompt will be detailed. I will give you the whole prompt and explain why it is written that way. I will also show you a trick to debug prompts better. Here is the prompt first: ```text from textwrap import dedent SYS_PROMPT = dedent(""" ## Role You are a university professor who is good at breaking down complex tasks into smaller parts that can be solved using Python code. ## Task 1. **Task Breakdown**: Break the user's request into several smaller steps, each suitable for solving with Python code. 2. **Code Generation**: Turn the current step into Python code. 3. **Code Execution**: Use tools to run the code and get the results. 4. **Iterative Progress**: Decide the next step based on the previous result, and repeat the process until you get the final answer to the user's request. ## Requirements - Plan and execute only one step at a time. Do not skip or combine steps. - Keep repeating the process until the task is fully completed. ## Output - Explain your thinking for each step. - Keep the structure clear. - Use a relaxed but authoritative tone. - Use emojis appropriately to make things friendlier. - Provide the final result. - If the result is an expression, solve it as a floating-point number. - Do not say "Task completed." ## Code Guidelines - The code runs in a Jupyter environment, and you can reuse variables that have already been declared. - Write code in an incremental way and use the kernel's statefulness to avoid repeating code. ## Python Package Management 1. You can only use numpy, pandas, sympy, scipy, and numexpr. 2. You are not allowed to install packages yourself using `pip install`. """) ``` I used markdown to organize the prompt. **Role** and **Output** parts set the tone and format for the answer. **Task** and **Requirements** tell the agent to plan only one step at a time in iterative style. **Code Guidelines** and **Python Package Management** set rules for writing Python and what third party libraries are allowed. One handy prompt debug trick is to write a `main` method in `agents.py` that asks the agent to repeat the instructions in detail. ```python async def main(): async with AOCAssistant() as agent: await Console(agent.run_stream(task=dedent(""" Please repeat my instructions in detail. """))) if __name__ == "__main__": asyncio.run(main()) ``` Then the agent will output its understanding of your instructions. ![The agent returns its understanding of my instructions.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/10/image-6.png) The agent returns its understanding of my instructions. Image by Author This helps find missing points and you can copy the agent’s version back into your prompt. ### Build the UI with chainlit I want my agent to be easy to use with a UI. Chainlit is a fast way to make a prototype. I put chainlit code in `app.py`. During development, you can run `chainlit run app.py -w` for hot reload. I first define `on_chat_start` and `on_chat_end` to initialize our custom Agent and manage the Jupyter server lifecycle. ```python @cl.on_chat_start async def on_chat_start(): assistant = AOCAssistant() await assistant.start() cl.user_session.set("assistant", assistant) @cl.on_chat_end async def on_chat_end(): assistant = cl.user_session.get("assistant") await assistant.stop() ``` In `on_message` we get user files, then call the agent and filter the returned text to show in the UI. ```python @cl.on_message async def on_message(message: cl.Message): input_msg = message.content file_path = None file_name = None if len(message.elements)>0: file_path = message.elements[0].path file_name = message.elements[0].name assistant: AOCAssistant = cl.user_session.get("assistant") output_msg = cl.Message(content='') async for event in assistant.run_stream( task=input_msg, file_name=file_name, file_path=file_path, ): if isinstance(event, ModelClientStreamingChunkEvent): await output_msg.stream_token(event.content) await output_msg.update() ``` And that is it. The agent app is done. It is simpler than it sounds. Agent app development is like this. Once you know what to do, coding is easy. --- ## After Class Practice Today, our project starts a Jupyter container with Docker SDK through Autogen. This method is fine for local testing. In enterprise apps, the agent itself will be in a container, so starting Jupyter with Docker SDK becomes hard. You need another way. To keep the code simple, we built a single-agent app. This makes the agent’s job more complex, which goes against the idea that each agent should have only one atomic job. In later optimization, you should try splitting into a multi-agent app. --- ## Course Summary Agent development tech will keep growing, and agents will handle more complex tasks. Code generation by agents for solving tasks will become more common. In today’s project, we followed ReAct agent design to build a single-agent app. It plans steps one by one and runs Python snippets in a stateful runtime container. This makes the agent smarter and more independent. I tested it with the 2024 Advent of Code challenge and got good results. The agent can also work in other complex scenarios. Thanks for reading today’s post. I am collecting ideas for agent development. If you have thoughts, share them in the comments. I will reply soon. --- ## Reference Answer The source code for this project is stored here. Feel free to use it. [agentic-ai-playground/07\_Complete\_Advent\_of\_Code at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-9.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-7)](https://github.com/qtalen/agentic-ai-playground/tree/main/07%5FComplete%5FAdvent%5Fof%5FCode?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### I Used Autogen GraphFlow and Qwen3 Coder to Solve Math Problems — And It Worked URL: https://www.dataleadsfuture.com/i-used-autogen-graphflow-and-qwen3-coder-to-solve-math-problems-and-it-worked/ Last updated: 2026-07-28T01:39:12.000Z *Disclaimer: This post contains affiliate links. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* Updated July 28, 2026: There's a concept that's been making waves again lately called Graph Engineering. It's built around three core ideas: 1. Breaking down a complete task into smaller units of work that become nodes in the graph. These nodes can be an agent, a direct LLM call, a function call, or even a human feedback step. 2. The edges between nodes determine the order in which they run. They can be sequential, looping, or parallel (Fan out and Fan in). 3. The edges between nodes determine the order in which they run. They can be sequential, looping, or parallel (fan-out and fan-in). Here's the thing though: Graph Engineering isn't actually a new concept. Way back in 2024, people already figured out that using DAGs to define workflows could handle complex tasks really well. That's exactly what gave birth to two frameworks: LangGraph and Autogen's GraphFlow. ![From Loop Engineering(ReAct) to Graph Engineering.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2026/07/opencode_loop-graph.drawio-2.png) From Loop Engineering(ReAct) to Graph Engineering. Image by Author --- In this tutorial today, I will show you how I used Autogen’s latest [GraphFlow](https://microsoft.github.io/autogen/stable/user-guide/agentchat-user-guide/graph-flow.html?ref=dataleadsfuture.com) with the Qwen3 Coder model to accurately solve all kinds of math problems — from elementary school level to advanced college math. As always, I put the source code for this tutorial at the end of the article. You can read it anytime. --- ## Course Introduction You must have secretly thought about using an LLM to do your homework. I had that idea from the very first day LLMs appeared, let them solve math problems for me. But dreams are sweet, reality is tough. If I throw a math problem at an LLM, it either gets the answer wrong or fails to reason at all. Here is what happened when I gave a math problem to the latest DeepSeek V3.1: ![A simple math problem took deepseek 580 seconds to finally solve.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/image.png) A simple math problem took deepseek 580 seconds to finally solve. Image by Author Not really the LLM’s fault. By design, LLMs are token generation models. Even if they can solve math problems, it is based on patterns learned during training, not actual numerical computation. No worries. Even if we cannot ask LLMs to solve math problems directly, we can ask them to generate code that solves the problem. That is something LLMs do very well. For example, I built a multi-agent app using Autogen GraphFlow and Qwen3 Coder that can accurately solve math problems from elementary school to university level: ![Easily solve the same problem with a multi-agent app based on code generation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/qwen3_coder.gif) Easily solve the same problem with a multi-agent app based on code generation. Image by Author The trick is simple: Qwen3 Coder writes Python code to solve the problem. A Docker Executor runs the code and gets the result. Then the app writes the final answer. Fun right? Want to learn how? Don’t rush. In this lesson, I will walk you through building this multi-agent app step by step. ### What Will I Learn? In this lesson, you will learn: 1. How to set up a Docker Executor agent that can run Python code. 2. How to build an Autogen GraphFlow workflow using multiple atomic agent nodes to work together on complex tasks. 3. Using a reasoning agent to help the workflow plan how to solve hard problems. 4. How to use a reflection agent to let the workflow review and improve its own code and results. The best part is I used small models like `qwen3-30b` and `qwen3-coder-plus`. And The final result beats large language models like DeepSeek v3.1 in numerical problem solving. ### Atomic Capability Agents In enterprise applications, the latest trend is using multiple atomic agents, each does one simple task. When you arrange them cleverly, they can handle complex jobs. This gives you two big wins: 1. Each agent has one clear job. That means fewer mistakes and less hallucination. 2. You don’t need to write long, confusing prompts. Just tell the agent what to do in simple words. In this tutorial, I will show you how this approach gives your enterprise AI applications a real edge. Ready? Buckle up. Let’s go. --- ## Pre-Class Prep Before we start, we need to get the environment ready. Yes, the Python code generated by Qwen can run in your local virtual environment. But to keep your system safe from bad code, I recommend using Docker to create a dedicated runtime. There are many guides on installing Docker, so I will skip that part. Today, I will just show you how to prepare the Docker image. Here is the `Dockerfile`, used for the container in this tutorial: ```text FROM python:3.13-slim-bookworm WORKDIR /app ENV PYTHONDONTWRITEBYTECODE=1 \ PYTHONUNBUFFERED=1 \ PIP_NO_CACHE_DIR=1 COPY requirements_docker.txt requirements.txt RUN pip install --no-cache-dir --upgrade pip && \ pip install --no-cache-dir -r requirements.txt --upgrade ``` One special thing: we need to install some math packages in this Python runtime. That helps the agent solve all kinds of math problems better. I put all dependencies into `requirements.txt`: ```text numpy==2.3.0 pandas==2.3.2 sympy==1.12 scipy==1.16.1 ``` You probably already thought of this: if I can install math packages, I can also install packages for other subjects. Just say the word. You need to build the Python runtime container into an image ahead of time: ```shell docker build -t python-docker-env . ``` Also, to make the Docker Executor work, you need to install Autogen’s Docker extension into your project’s virtual environment. ```shell pip install -U autogen-ext[docker] ``` Unlike other workflow frameworks, Autogen GraphFlow can filter messages. It lets you choose which messages go to the next node. That saves tokens and reduces hallucination. If you want to trace message inputs and outputs at each node, you can use MLFlow. I wrote a whole article about that. We will use it soon. [Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracing TutorialEnhance your multi-agent application’s observability, explainability and Traceability![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-39.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/generated_00--2-.webp)](https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/) Now let’s start the lesson. --- ## Ding Ding — Class Begins ### Set Up the LLM Clients In today’s multi-agent application, we will use two Qwen3 models. The main one is `qwen3-coder-plus`. I will use this model to generate Python code that solves problems. The other model is `qwen3-30b-a3b-instruct-2507`, It is a MoE model — small but smart. I use it in the workflow for thinking and code review. This model cannot do deep reasoning, but it saves us lots of tokens and time. I can also build a separate reasoning agent to handle problem-solving logic. This setup is flexible and works well. I will prove that to you soon. To make Autogen work with Qwen3 models, I prepared an OpenAILike client for you. You can click here to learn more. [Build AutoGen Agents with Qwen3: Structured Output & Thinking ModeSave yourself 40 hours of trial and error![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-38.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Autogen_Qwen3_cover_2-2.webp)](https://www.dataleadsfuture.com/build-autogen-agents-with-qwen3-structured-output-thinking-mode/) I want the `qwen3-30b` model to think through solutions more carefully, so I lowered the temperature to 0.1: ```python slm_client = OpenAILikeChatCompletionClient( model="qwen3-30b-a3b-instruct-2507", temperature=0.1 ) ``` I want the `qwen3-coder` model to be precise and consistent. So I set its temperature to 0.01: ```python coder_client = OpenAILikeChatCompletionClient( model="qwen3-coder-plus", temperature=0.01 ) ``` ### Prepare the Docker Executor Autogen’s Code Executor comes in two types. `CommandLineCodeExecutor`, runs Python code once then exits. `JupyterExecutor` keeps state between multiple code runs. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-2.two_executor.drawio.png) The differences between the two kinds of Code Executor. Image by Author For today’s task, `DockerCommandLineCodeExecutor` is enough. But after many real project tests, a stateful Python runtime based on Jupyter is much better for agents to solve complex tasks by exploring on their own: [Exclusive Reveal: Code Sandbox Tech Behind Manus and Claude Agent SkillsUse Jupyter code executor to help your agent finish tasks in a smarter way![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-43.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4.webp)](https://www.dataleadsfuture.com/exclusive-reveal-code-sandbox-tech-behind-manus-and-claude-agent-skills/) We will build it using the image we prepared earlier. Since some data science code takes time to run, I suggest you set a longer timeout: ```python docker_executor = DockerCommandLineCodeExecutor( image="python-docker-env", timeout=300, ) ``` The Docker Executor container removes itself after the agent finishes. If you want to keep it for debugging, set `auto_remove` to False. I am not sure if Autogen’s Docker Executor has a bug. If you want to mount your host file system into the container, use the `extra_volumes` parameter. Avoid the workspace folder inside the container. ```python docker_executor = DockerCommandLineCodeExecutor( image="python-docker-env", timeout=3000, extra_volumes={str(Path("./data").resolve()): {"bind": "/data", "mode": "rw"}}, ) ``` ### Write Your Agents To keep things clean, I put agents and their prompts in separate files: `agents.py` and `prompts.py`. This project has three core agents: `coder` agent to generate code `exe_agent` to run the code and `reviewer` agent to check code and results. ![The three core agents of this project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow_three_agents.drawio.png) The three core agents of this project. Image by Author The `coder` agent uses Autogen’s `AssistantAgent`. Its job is to understand the user’s question, plan a solution, and write executable Python code. It uses the `qwen3-coder-plus` model: ```python coder = AssistantAgent( "coder", model_client=coder_client, system_message=PROMPT_CODER ) ``` Here is its prompt: ```python PROMPT_CODER = dedent(""" ## Role You are a developer engineer responsible for writing Python code. ## Task For the user's question, you will generate a piece of **Python** code that runs correctly and gives the right result. ### Requirements - **Write the code logic strictly following thinker's problem-solving approach.** - Use the print function to show the result. - Short code comments. - Only output the code block, no extra words or explanations. - When showing results, if it's a float, keep two decimal places. ### Available libraries - numpy, pandas, sympy, numexpr, scipy """) ``` I used a popular structured prompt method. Markdown syntax and simple language make the prompt clean and easy — almost like a programming language. I also told the agent which third-party packages it can use. This is important. Knowing what tools are available helps the agent do better work and avoid mistakes. The `exe_agent` is an instance of `CodeExecutorAgent`. It is just a wrapper around the Docker Executor. Simple: ```python exe_agent = CodeExecutorAgent( "exe_agent", code_executor=docker_executor, ) ``` The `reviewer` agent is also an `AssistantAgent` but uses the `qwen3-30b` model. Its job is to check the code and results from the coder. If it passes, the `reviewer` writes the answer. If it fails, the `reviewer` gives feedback: ```python reviewer = AssistantAgent( "reviewer", model_client=slm_client, system_message=PROMPT_REVIEWER, ) ``` Here is the reviewer’s prompt: ```python PROMPT_REVIEWER = dedent(""" ## Role You are a test engineer. ## Task - Check if the Python code follows the problem-solving approach proposed by thinker. - Verify if the code runs correctly and produces results. - Provide the review results. ### Review Passed "COOL" ### Review Failed - "REJECT" - Output the exact error message from exe_agent. - Give brief suggestions for improving the code. - Don't include any introductions or explanations. - Do not provide the revised code. """) ``` Again, I used simple structured text. I asked the agent to output different keywords for pass or fail. There is a problem here: Earlier, I said we want atomic agents that each do one job. But `coder` and `reviewer` are clearly doing multiple things. That is because I wanted to test if using LLMs to generate code for math problems even works. Once that is proven, I will refactor each agent’s prompt to make their jobs more focused. ### Build a Quick Prototype to Test the Idea Time to build a quick prototype to show the boss that this idea works. In the prototype stage, I don’t want to work late building a complex Autogen GraphFlow. I want to solve it in one line of code. In Autogen, the simplest way to run agents in a loop is `RoundRobinGroupChat`. `RoundRobinGroupChat` is an infinite loop of agents. It only stops when it hits a termination condition. ![RoundRobinGroupChat is an infinite loop of agents.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-3.round_robin.drawio.png) RoundRobinGroupChat is an infinite loop of agents. Image by Author I set a condition to stop when the message contains “COOL”. Then I put all agents into the `RoundRobinGroupChat` in order. ```python text_termination = TextMentionTermination("COOL") team = RoundRobinGroupChat([coder, exe_agent, reviewer], termination_condition=text_termination) ``` Let’s test it with a math problem: ```python async def main(): task = "The bag has 4 black balls and 1 white ball. Every time, you pick one ball at random and swap it for a black ball. Keep going, and figure out the chance of pulling a black ball on the third try." async with docker_executor: await Console(team.run_stream(task=task)) asyncio.run(main()) ``` One thing to remember: when you run agents with Docker Executor, you need to manage the container’s start and stop. The best way is to use `async with`. ![The project prototype built using RoundRobinGroupChat seems pretty good.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/round_robin_group_chat-1.gif) The project prototype built using RoundRobinGroupChat seems pretty good. Image by Author Looks good. The agent understood the question, wrote Python code, and got the result. The `reviewer` checked it and wrote the answer. ### Spot Problems in the Prototype and Plan Improvements As a quick prototype, `RoundRobinGroupChat` is enough to prove the idea works. But if you put this code into production, you will find it unstable and inaccurate. Before turning this into production code, we need to identify what is wrong and how to fix it. Try the prototype a few times. For simple problems it works. For more challenging math problems, the coder writes incorrect code. Why? Because the `coder` has to both plan the solution and write the code. For hard problems, it cannot think of a good plan, so the code logic fails. The `reviewer` has the same issue. It checks the code, writes the answer if it passes, or gives feedback if it fails. It is too busy. It mixes all jobs. The answer often includes code comments. That is hard for students who just want to learn how to solve the problem, not read Python code. Remember our agent design rule? In a multi-agent system, each agent should do one atomic job. That makes the whole system perform better. So our next step is clear. We need to split the jobs of `coder` and `reviewer`. Let `coder` only write code. Let `reviewer` only check if the code passes or fails. We also need two new agents: ![The responsibilities of the other two agents.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-4.two_more.drawio-1.png) The responsibilities of the other two agents. Image by Author The `thinker` agent acts as the brain of the system. It uses CoT thinking to plan the solution. The `coder` follows the `thinker’s` plan exactly — no more thinking. The `writer` agent writes the solution and answer in student-friendly language based on the correct code and result. The `reviewer` no longer writes answers, it only checks code. Here is the new workflow: ![A diagram showing a project created using Autogen GraphFlow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-5.workflow.drawio.png) A diagram showing a project created using Autogen GraphFlow. Image by Author As you can see, `thinker` and `writer` are outside the loop. `RoundRobinGroupChat` won’t work anymore. We need a more complex workflow. Ladies and gentlemen, please welcome Autogen GraphFlow. ### Prepare the New Agents Before coding the `GraphFlow`, we need to add two new agents as planned. First, the `thinker` agent. Even though it is the brain, it still uses the `qwen3-30b`, model. We only need it to plan the solution: ```python thinker = AssistantAgent( "thinker", model_client=slm_client, system_message=PROMPT_THINKER, ) ``` Its prompt is simple and structured: ```python PROMPT_THINKER = dedent(""" ## Role You are a college professor who is good at breaking down complex problems into clear steps for solving them. ## Task For the user's question, break it down into a problem-solving approach that can be calculated step by step using code. ## Response Use an ordered list to show the steps. ### Requirements - Do not include actual code. - Do not solve the problem. - Do not do any numerical calculations. """) ``` Second, the `writer` agent. It also uses the `qwen3-30b`. It writes the answer based on correct code and results. I will show you later how to pick the correct code: ```python writer = AssistantAgent( "writer", model_client=slm_client, system_message=PROMPT_WRITER, model_client_stream=True ) ``` Here is the `writer’s` prompt: ```python PROMPT_WRITER=dedent(""" ## Role You're a college professor who's really good at explaining problem-solving ideas and answers in a way students can easily understand. ## Task Based on the user's question, use the logic of the Python code and the result of code execution to write the answer. ## Answer Content The answer should include the problem-solving idea and the final answer. ### Style - Use natural language that humans can understand. - **Don't include any code**. ### Note - **Use the execution result from exe_agent as the answer**. - The problem-solving idea must match the logic in the coder's code. - You can't come up with your own problem-solving idea. """) ``` With both agents ready, we can now build the final agent workflow using Autogen GraphFlow. ### What Is Autogen GraphFlow When people talk about Autogen, they say its biggest difference from LangGraph or CrewAI is that Autogen lets LLMs choose which agent to use to plan and complete tasks. So Autogen feels more suited for research. But for enterprise applications, we often want consistent and robust code execution. So we prefer to pre-arrange how tasks flow between agents using workflows. Autogen’s team saw this too. In a recent version, they released `GraphFlow`— their workflow coding framework. Compared to other workflow frameworks, `GraphFlow` has two special features: It can filter messages going into a node. Only the messages the LLM needs to know are kept. That reduces hallucination. ![GraphFlow can filter messages going into a node.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-7.Filter.drawio.png) GraphFlow can filter messages going into a node. Image by Author It can group edges going into a node. You can set the node to run after all edges finish or just after some edges finish. That lets you build complex branches or loops. ![Two types of Activation Condition in the edges of GraphFlow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-8.activation_node.drawio.png) Two types of Activation Condition in the edges of GraphFlow. Image by Author In today’s application, we need both features. When code fails, the `reviewer` sends feedback to the `coder`. When code passes, the `writer` only reads the correct code, no noise from wrong code. In the end, we'll have a multi-agent app with a chat interface, and it can perfectly answer your math questions: ![The beautiful user interface of our multi-agent application.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/chainlit_app-1.gif) The beautiful user interface of our multi-agent application. Image by Author Let’s begin. ### Message Filtering and Edge Grouping All agents are ready. Now we just need to focus on building the `GraphFlow`. Here is a diagram of what the final workflow should look like: ![The diagram of the final workflow should look like.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-5.workflow.drawio-1.png) The diagram of the final workflow should look like. Image by Author You can see a loop from `coder` to `reviewer`. The `reviewer` keeps checking the code and uses two keywords to decide the next step. The loop may run many rounds and create many messages. But the `reviewer` only needs the latest Python code and the result. So we need a `MessageFilterAgent` to filter messages: ```python filtered_reviewer = MessageFilterAgent( "reviewer", wrapped_agent=reviewer, filter=MessageFilterConfig( per_source=[ PerSourceFilter(source="user", position="first", count=1), PerSourceFilter(source="thinker", position="first", count=1), PerSourceFilter(source="coder", position="last", count=1), PerSourceFilter(source="exe_agent", position="last", count=1), ] ) ) ``` This agent wraps the `reviewer` and filters messages going into it. Messages from `user` and `thinker` take the first one — they are not in the loop so the earliest message is fine. Messages from `coder` and `exe_agent` take the latest one — as the loop runs they create many messages but only the last one matters. The `writer` agent also needs a `MessageFilterAgent`. It needs the final correct code and result to write the answer: ```python filtered_writer = MessageFilterAgent( "writer", wrapped_agent=writer, filter=MessageFilterConfig( per_source=[ PerSourceFilter(source="user", position="first", count=1), PerSourceFilter(source="coder", position="last", count=1), PerSourceFilter(source="exe_agent", position="last", count=1), ] ) ) ``` After creating both `MessageFilterAgents` let’s look at another issue: The diagram shows two edges going into the `coder` node. Should the `coder` wait for both edges to finish or just one? ![Should the coder wait for both edges to finish or just one?](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/image-4.png) Should the coder wait for both edges to finish or just one? Image by Author You might say of course just one. We know that but the code does not. Also different graph-based Workflow frameworks handle multi-edge nodes differently. For example, LlamaIndex Workflow runs when any edge arrives. But Autogen Graphflow by default waits for all edges to finish. So in the latest Autogen version, `GraphFlow` added a new feature called `activation_group`. This groups edges going into a node. With `activation_condition`, you can set if all edges must finish or just one. In today’s project, we put both edges into the same group and set `activation_condition` to `any` — meaning run when any edge arrives. ### Start Coding the Workflow We know all agents. We know how to build the `GraphFlow`. Now coding the workflow is easy. Since the workflow has a loop, we need a guardrail to stop it from running forever. I use `MaxMessageTermination` to stop when the message count hits a limit. ```python terminator = MaxMessageTermination(20) ``` Next create a `DiGraphBuilder` and add nodes. For `reviewer` and `writer`, use the filtered versions — `filtered_reviewer` and `filtered_writer`. ```python builder = DiGraphBuilder() builder.add_node(thinker) builder.add_node(coder).add_node(exe_agent) builder.add_node(filtered_reviewer).add_node(filtered_writer) builder.set_entry_point(thinker) ``` If your first node is inside a loop, set it as the entry. Here `thinker`, is not in the loop but I still set it as an entry. Then add edges between nodes. As we said, set `activation_group` for the two edges going into `coder` as `gen_code` and `activation_condition` as `any`. ```python builder.add_edge(thinker, coder, activation_group="gen_code", activation_condition="any") builder.add_edge(coder, exe_agent) builder.add_edge(exe_agent, filtered_reviewer) builder.add_edge(filtered_reviewer, filtered_writer, condition="COOL") builder.add_edge(filtered_reviewer, coder, condition="REJECT", activation_group="gen_code", activation_condition="any") ``` Set different keywords as conditions for edges from `filtered_reviewer`, to different nodes based on code review results. Finally, put all parameters together and initialize the `GraphFlow` instance: ```python graph = builder.build() self._flow = GraphFlow(builder.get_participants(), graph=graph, termination_condition=terminator) ``` To hide details and make it easy for users to call this workflow, we can wrap the code in a `SuperTeacherFlow` Class. To work with Autogen `Console`, we also add `run` and `run_stream` methods: ```python class SuperTeacherFlow: def __init__(self): self._build_workflow() async def run( self, *, task: str | BaseChatMessage | Sequence[BaseChatMessage] | None = None, cancellation_token: CancellationToken | None = None, output_task_messages: bool = True, ) -> TaskResult: result: TaskResult | None = None async for message in self.run_stream( task=task, cancellation_token=cancellation_token, output_task_messages=output_task_messages, ): if isinstance(message, TaskResult): result = message if result is not None: return result raise AssertionError("The stream should have returned the final result.") async def run_stream( self, *, task: str | BaseChatMessage | Sequence[BaseChatMessage] | None = None, cancellation_token: CancellationToken | None = None, output_task_messages: bool = True, ) -> AsyncGenerator[BaseAgentEvent | BaseChatMessage | TaskResult, None]: async for event in self._flow.run_stream( task=task, cancellation_token=cancellation_token, output_task_messages=output_task_messages ): yield event def _build_workflow(self): ... ``` ### Test the GraphFlow Execution Before moving on, let’s test the workflow. While testing we can add MLFlow to trace LLM calls: ```python import mlflow mlflow.set_experiment("Super Teacher Workflow") mlflow.autogen.autolog() ``` From the animation, you can see that the workflow follows the preset nodes. `Thinker` plans the solution. `Coder` writes code based on the plan. `Reviewer` checks the code. `Writer` outputs the report. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/qwen3_coder-1.gif) The workflow in GraphFlow version runs really well. Image by Author With MLFlow, we can also check if the message filters work: ![Even though the code is generated many times, only the last generated code and result go into the writer node.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/image-2.png) Even though the code is generated many times, only the last generated code and result go into the writer node. Image by Author You can see that although `coder` generated code many times, only the last code and result went to the `writer` to write the report. Our `GraphFlow` is built right. ### Build a User Interface with Chainlit Maybe your end user is just a student who needs help with homework and knows no Python. Then you need a GUI. I chose chainlit to build a simple chat interface. Before coding, we need to adjust `SuperTeacherFlow`: In the earlier code, we had to use `async with` to manage the `docker_executor` container start and stop every time. I could put the context manager inside `SuperTeacherFlow’s` `run_stream` method — start container when solving starts, stop when done. But that makes the container start and stop too often, it's a waste of resources. So I put the container start and stop in two separate methods. Start when the user opens the chat page. Stop when the user closes it: ```python class SuperTeacherFlow: def __init__(self): self._build_workflow() ... @staticmethod async def start(): await docker_executor.start() @staticmethod async def stop(): await docker_executor.stop() ... ``` These match Chainlit’s lifecycle events `on_chat_start` and `on_chat_end`: ```python @cl.on_chat_start async def on_chat_start(): flow = SuperTeacherFlow() await flow.start() cl.user_session.set("workflow", flow) @cl.on_chat_end async def on_chat_end(): flow = cl.user_session.get("workflow") await flow.stop() ``` Since users do not need to see internal steps, I hide outputs from internal nodes. I just show mask text to show progress: ```python @cl.on_message async def main(message: cl.Message): workflow = cl.user_session.get("workflow") show_text = "Thinking..." msg = cl.Message(content=show_text) await msg.send() streaming = False async for event in workflow.run_stream(task=message.content): if isinstance(event, TaskResult): continue if event.source and event.source in MASK_MAPPING.keys(): show_text = MASK_MAPPING[event.source] msg.content = show_text await msg.update() if isinstance(event, ModelClientStreamingChunkEvent): if not streaming: show_text = "" msg.content = show_text await msg.update() streaming = True await msg.stream_token(event.content) await msg.update() ``` Now let’s see the final result: ![The final interface effect of the project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/chainlit_app.gif) The final interface effect of the project. Image by Author See? Ask any math question. The agent runs code, then answers in natural language with solution steps. Our project is a full success! --- ## Homework Here are ways to improve our project. I will not implement them here due to space. Your turn to practice. ### Exercise One: Add Few Shot Examples to Thinker Node Even with atomic agents, sometimes hard problems still get wrong answers. Mistakes happen in two places: wrong Python code — fix by making `coder` follow thinker’s plan exactly. Or a wrong solution plan from `thinker`, which needs human fix. Here is a way to fix `thinker’s` wrong plan and save the knowledge: Earlier, the solution plan relied only on the LLM’s own ability. But we use a small model — thinking power is limited. When we find `thinker’s` plan is wrong, we can write a correct plan and store it with the question in a vector database. Next time for similar questions, `thinker` (or a new node) can use Agentic RAG to find the correct plan. `Thinker` will follow that plan. Accuracy will go up. Here is the process: ![Add solution ideas written by humans to the Workflow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/09/Super_Teacher_Flow-Page-10.human_in_the_loop.drawio.png) Add solution ideas written by humans to the Workflow. Image by Author ### Exercise Two: Add OCR and Drawing Ability Now we type questions by hand. But if your users are students, you need OCR to read questions from textbooks — especially ones with lots of formulas. In Autogen, you can use `MultiModalMessage` to send images. But the current version seems to have a bug, it cannot send image URLs or base64 data correctly. So you either fix `MultiModalMessage` or use OpenAIClient directly. On the output side, since we help students from elementary to high school, diagrams are more fun than plain text. Good news — many new drawing models handle text well. I think rendering the solution as a diagram is worth a try. --- ## Lesson Wrap-Up When LLMs first came out, everyone threw math problems at them and said LLMs are bad at math. Even though new models got better at numbers, they still struggle with hard math. Then came the agent era. We let LLMs write code, then run the code to solve problems. Now we can solve all kinds of math problems well. If we let LLMs use third-party packages in code, we can solve way more than just numbers. In this tutorial, I combined Autogen GraphFlow and Qwen3-coder to show you how I built a multi-agent application that uses LLMs to write code for math problems. This application proves that multiple atomic agents — each focused on one job — can work together to match or beat big LLMs. Less hallucination, better instruction following. By adding reasoning and reflection agents, we boost the whole app’s performance. Message filters and guardrails keep the workflow running as planned. This tutorial covered the general process and core ideas of enterprise agent development. I could not cover everything. If you have questions, leave a comment. I am collecting ideas for agent apps. If you want to build something with LLMs, leave me a note. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. No spam, No ads, you can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Source Code Get the source code for this article. ****Start using it right away and save yourself a ton of trial and error time.** [Grab the Source Code ](#/portal/signup) ### Use LLamaIndex Workflow to Create an Ink Painting Style Image Generation Workflow URL: https://www.dataleadsfuture.com/use-llamaindex-workflow-to-create-an-ink-painting-style-image-generation-workflow/ Last updated: 2026-04-22T08:24:36.000Z In today's article, I'll help you build a workflow that can generate ink illustrations with strong Eastern style. This workflow also allows for multiple rounds of prompt adjustments and final image tweaks, helping save on token and time costs. You can find the source code for this project at the end of the article. --- ## Introduction Recently I wanted to create an agent workflow that could quickly generate images for my blog at low cost. I wanted my blog images to have strong artistic flair and classical Eastern charm. So I hoped my workflow could precisely control the LLM context and continuously adjust the prompts for drawing as well as the final image effects, while keeping token and time costs to a minimum. Then I immediately faced a dilemma: If I chose low-code platforms like dify or n8n, I wouldn't get enough flexibility. These platforms can't support adjusting prompts in conversations or generating blog images according to article styles. If I chose popular agent development frameworks like LangGraph or CrewAI, these frameworks are too high-level in abstraction, preventing fine control over the execution process of agent applications. If you were faced with this task, how would you choose? Fortunately, the world isn't black and white. After countless failures and continuous attempts, I finally found a great solution: LLamaIndex Workflow. It provides an efficient workflow development process while not abstracting too much from my agent execution process, allowing the image generation program to run precisely as I require. Today, let me use the latest [LlamaIndex Workflow 1.0](https://www.llamaindex.ai/blog/announcing-workflows-1-0-a-lightweight-framework-for-agentic-systems?ref=dataleadsfuture.com) version to build a workflow for generating ink painting style illustrations for you. ### Why should you care? In today's article, I will: 1. Guide you to learn the basic usage of LlamaIndex Workflow through project practice. 2. Use chainlit to build a chatbot interface where you can visually see the generated images. 3. Leverage deepseek to generate more project-appropriate drawing prompts outside of the DALL-E-3 model. 4. Use multi-turn conversations to make further adjustments to the generated prompts or the final generated images. 5. Optimize costs at the token and time level through fine control of the LLM context. Ultimately, you only need a simple description to draw a beautiful ink painting style image. ![Use a workflow to create a beautiful ink-style illustration.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/n7mxRwCFZhWwzAMXAVE0bg1PzS8S6R.webp) Use a workflow to create a beautiful ink-style illustration. Image by DALL-E-3 More importantly, through practicing this project, you will gain a preliminary understanding of how we use workflows to complete complex customized requirement development in enterprise-level agent applications. --- If you need some prerequisite knowledge, I wrote an article explaining in detail the event-driven architecture of LlamaIndex Workflow. You can read it by clicking here: [Deep Dive into LlamaIndex Workflow: Event-driven LLM architectureWhat I think about the progress and shortcomings after practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-36.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LLamaIndex_Workflow-4.webp)](https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/) --- ## Business Process Design For the development of agent workflow type applications, I strongly recommend designing the business process flow before starting coding. This helps you grasp the entire program operation process. In this chapter, I will demonstrate my complete design thinking for the business process flow of this project: ### Prompt generation process In today's project, I'm using DALL-E-3 for drawing. DALL-E-3 itself has the ability to rewrite user intentions into detailed prompts suitable for drawing. But our requirements are higher. We want DALL-E-3 to draw the picture exactly as I imagine. So we will move the process of rewriting user intentions into detailed prompts from the DALL-E-3 model to our own workflow node. Since today's theme is to draw beautiful ink painting style illustrations, I need to select an LLM that can fully understand the Eastern ambiance in user intentions and expand it into a drawing prompt that DALL-E-3 can understand. Here I chose the DeepSeek-Chat model. Since DeepSeek has been prompted to generate DALL-E-3 drawing prompts, the generated prompts are in pure English. If you are proficient in English, then this step can be ended here. But if you, like me, are a non-native English speaker and want to accurately understand the content of the generated prompts, you can add a translation node. This step is not troublesome. Finally, the generated prompts and their translations will be returned to the user through LlamaIndex Workflow's StopEvent. ![The process of creating drawing prompts.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-24.png) The process of creating drawing prompts. Image by Author ### Using context sharing to decouple workflow loops After generating the prompt, next we either provide the prompt to DALL-E to generate the image or return to let DeepSeek readjust the prompt again. At this point, you must be thinking of adding user feedback and workflow iteration features. Based on the final generated image effect, provide modification suggestions and ask the workflow to regenerate the prompt, looping until satisfied output is obtained. Since we need to support user feedback after both prompt generation and image generation nodes, this loop will greatly increase the complexity of the workflow. But in today's project, I plan to use a small trick to significantly simplify the implementation of the workflow. Instead of adding user feedback after prompt or image generation to determine whether to regenerate the prompt, we add an if-else node at the very beginning of the workflow to judge whether the user input contains specific keywords (here APPROVE, you can replace it with your own). If it does contain, go to the image generation branch; if not, go to the prompt generation branch. This way, each iteration is a re-execution of the workflow, thus decoupling the workflow from the iteration. ![Use conditional branches and Context to separate workflow loops.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-25.png) Use conditional branches and Context to separate workflow loops. Image by Author After adding the branch node, we will face a problem: the image generation branch doesn't know what the prompt generated from the previous run was. And I don't want to save all the message history from the last run because the message history also includes translations of the prompt and other information that doesn't need to be sent to the LLM. At this point, the best choice is to let all runs of the workflow use the same context and save the generated prompt into the context. Fortunately, LlamaIndex Workflow supports sharing the same context across multiple runs, so we can add logic to save variables into the Workflow Context. ### Rewrite user's historical drawing requests Since users will gradually adjust the prompts generated by the LLM through multiple conversations, we need to provide the LLM with the complete conversation history. The common approach is to use messages with role as user and role as assistant to save and provide the historical conversations between users and the LLM to DeepSeek. But doing this, as the adjustments continue, the conversation messages will become longer and longer, causing the LLM to ignore truly important key information. We can't adopt the method of truncating historical information and only keeping the recent few rounds of conversations. Because the most detailed drawing intention is usually provided in the user's initial request. So here I will take the approach of rewriting the user's request history, rewriting the user's multiple rounds of adjustments to the prompt and image into a complete request. The specific method is: create a list container in the Context, when generating the prompt, append the user's latest input into this list. Then call DeepSeek to rewrite all user inputs into a complete drawing description and store it in the Context. ![We need to rewrite the user's past input fragments into one complete drawing intention. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-26.png) We need to rewrite the user's past input fragments into one complete drawing intention. Image by Author ### Improve system prompt Finally, we modify the node where DeepSeek generates drawing prompts. Before instructing the LLM to generate prompts, retrieve the rewritten user historical requests and the last generated drawing prompt from the Context, and merge them into the system prompt. At the very first execution of the workflow, the user's historical requests and the last generated prompt in the Context are empty. Still, it doesn't matter because the latest user requests are always provided to the LLM with role as user message. Thus, the complete business process diagram is designed, as shown in the figure below. ![Complete business process diagram.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-27.png) Complete business process diagram. Image by Author Next, we can start coding according to the design of the business process diagram. --- ## Develop Your Drawing Workflow Using LlamaIndex Workflow In today's project, in addition to using LlamaIndex Workflow to build the drawing workflow, I will also utilize Chainlit to create an interactive interface, facilitating interaction with the workflow application. The dialogue interface is shown in the figure below: ![An interactive interface built with Chainlit.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-28.png) An interactive interface built with Chainlit. Image by Author Next, let's start writing the actual code logic. ### Overall project structure The code structure of this project is as follows: ```text project_source_code |-app.py |-ctx_manager.py |-events.py |-prompts.py |-workflow.py ``` `app.py` is the entry file of the application, used to store Chainlit code. `workflow.py` contains the core code logic implementation of LlamaIndex Workflow. `prompts.py` stores all prompts that will be provided to the LLM. `events.py` contains LlamaIndex Workflow related event definitions. `ctx_manager.py` stores all operations logic for Workflow Context. ### Environment variables In this project, since we need to use two sets of LLMs, DeepSeek and DALL-E-3, we need to prepare two sets of environment variables in the `.env` file: ```bash OPENAI_API_KEY= OPENAI_API_BASE= REAL_OPENAI_API_KEY= REAL_OPENAI_BASE_URL= ``` According to my habits, `OPENAI_API_KEY` and `OPENAI_API_BASE` still point to the DeepSeek service. And I use `REAL_OPENAI_API_KEY` and `REAL_OPENAI_BASE_URL` to point to OpenAI's service. You can adjust according to your own habits. ### Define several events needed by workflow LlamaIndex Workflow is based on an event-driven architecture, so transitions between code nodes need to define corresponding events. These event definitions are all placed in events.py. But before writing the core code of the Workflow, directly explaining Event definitions might make you feel a bit lost. Don't worry, because I've already explained the functions of various workflow nodes in detail earlier, so you can still understand where each event is used. `GenPromptEvent`, this event will drive the DeepSeek code node to start generating drawing prompts. The content attribute stores the user's latest input. ```python class GenPromptEvent(Event): content: str ``` `PromptGeneratedEvent`, when the drawing prompt is generated, this event will be thrown. The content contains the prompt generated by DeepSeek. Downstream nodes can subscribe to this event for prompt translation, image generation, and user request rewriting work. ```python class PromptGeneratedEvent(Event): content: str ``` `StreamEvent`, since I want to display the prompt content on chainlit through streaming output, I need a `StreamEvent`. The chainlit code can iterate over this event to get streaming messages. ```python class StreamEvent(Event): target: str delta: str ``` `GenImageEvent`, this event will drive the DALL-E-3 node to start generating images. The content attribute still contains the prompt generated by DeepSeek. ```python class GenImageEvent(Event): content: str ``` `RewriteQueryEvent`, this Event will call the LLM node to rewrite the user's historical input into a complete drawing intention. Since historical input is taken from Context, this event has no attributes. ```python class RewriteQueryEvent(Event): pass ``` ### Implement workflow code logic After defining all the Workflow events, next we can implement the Workflow code according to the previously drawn business process diagram. In `workflow.py`, we define a class named `ImageGeneration` as a subclass of LlamaIndex Workflow. In the `__init__` method, we need to initialize two LLM clients. One is used to generate drawing prompts with the DeepSeek model, using LlamaIndex's OpenAILike client. The other is the DALL-E-3 model, using OpenAI's client directly. ```python class ImageGeneration(Workflow): def __init__(self, *args, **kwargs): self.deepseek_client = OpenAILike( model="deepseek-chat", is_chat_model=True, is_function_calling_model=True ) self.openai_client = AsyncOpenAI( api_key=os.getenv("REAL_OPENAI_API_KEY"), base_url=os.getenv("REAL_OPENAI_BASE_URL"), ) super().__init__(*args, **kwargs) ``` The `on_start` method is the entry method of the workflow, it only does conditional branch judgment. If the user input contains "APPROVE", it throws the `GenImageEvent` to start drawing pictures. Otherwise, it throws the `GenPromptEvent` to start generating drawing prompts or adjust existing prompts. ```python class ImageGeneration(Workflow): ... @step async def on_start(self, ctx: Context, ev: StartEvent) -> GenImageEvent | GenPromptEvent: query = ev.query if len(query) > 0 and ("APPROVE" in query.upper()): return GenImageEvent(content=query) else: return GenPromptEvent(content=ev.query) ``` The `prompt_generator` method subscribes to the `GenPromptEvent` event, used to generate or adjust drawing prompts. ```python class ImageGeneration(Workflow): ... @step async def prompt_generator(self, ctx: Context, ev: GenPromptEvent) \ -> PromptGeneratedEvent | RewriteQueryEvent | None: user_query = ev.content hist_query = await ctx_mgr.get_rewritten_hist(ctx) hist_prompt = await ctx_mgr.get_image_prompt(ctx) system_prompt = PROMPT_GENERATE_SYSTEM.format( hist_query=hist_query, hist_prompt=hist_prompt ) messages = [ ChatMessage(role="system", content=system_prompt), ChatMessage(role="user", content=user_query) ] image_prompt = "" events = await self.deepseek_client.astream_chat(messages) async for event in events: ctx.write_event_to_stream(StreamEvent(target="prompt", delta=event.delta)) image_prompt += event.delta await ctx_mgr.add_query_hist(ctx, user_query) await ctx_mgr.set_image_prompt(ctx, image_prompt) ctx.send_event(PromptGeneratedEvent(content=image_prompt)) ctx.send_event(RewriteQueryEvent()) ``` The `prompt_generator` method first gets the rewritten user input historical intentions and the prompt generated in the last round of workflow from the Context, and integrates them with the predefined system prompt template. After integration, the system prompt will be submitted to the DeepSeek model together with the user's latest input to generate the latest drawing prompt. I called the llm client's streaming api, so I use StreamEvent to throw the messages returned by the large model into the stream of Context. At the same time, I will piece together all messages into a complete prompt and pass it down. Of course, after the prompt is generated, I will write back the user's latest input and the generated prompt into Context. The `translate_prompt` method is optional. Its function is to translate the prompts generated by DeepSeek so that you can accurately understand the content of the prompt. The translated results will only be displayed on the interface, so they won't be written into Context. ```python class ImageGeneration(Workflow): ... @step async def translate_prompt(self, ctx: Context, ev: PromptGeneratedEvent) -> StopEvent: image_prompt = ev.content messages = [ ChatMessage(role="system", content=PROMPT_TRANSLATE_SYSTEM), ChatMessage(role="user", content=image_prompt) ] events = await self.deepseek_client.astream_chat(messages) translate_result = "" async for event in events: ctx.write_event_to_stream(StreamEvent(target="translate", delta=event.delta)) translate_result += event.delta return StopEvent(target="prompt", result=translate_result) ``` The `translate_prompt` method will return `StopEvent`, marking the end of this workflow execution. If you don't need to translate the drawing prompt, you can return `StopEvent` directly in the `prompt_generate` method to end the workflow. If you include the keyword "APPROVE" in your input, the workflow will directly enter the `generate_image` method. This method will get the latest drawing prompt from Context and call the `_image_generate` method to start generating images. ```python class ImageGeneration(Workflow): ... @step async def generate_image(self, ctx: Context, ev: GenImageEvent) -> StopEvent: prompt = await ctx_mgr.get_image_prompt(ctx) image_url, revised_prompt = await self._image_generate(prompt=prompt) return StopEvent(target="image", result={ "image_url": image_url, "revised_prompt": revised_prompt } ) ``` We have already used DeepSeek to generate good drawing prompts in advance, but DALL-E-3 will still rewrite the input prompt out of safety considerations. This may cause the drawn image to deviate too far from our intention. So we can add the following content in front of the drawing prompt to prevent DALL-E-3 from rewriting the prompt: ```python class ImageGeneration(Workflow): ... async def _image_generate(self, prompt: str) -> tuple[str, str]: ## Stop DALL-E 3 from rewriting incoming prompts final_prompt = f""" I NEED to test how the tool works with extremely simple prompts. DO NOT add any detail, just use it AS-IS: {prompt} """ response = await self.openai_client.images.generate( model="dall-e-3", prompt=final_prompt, n=1, size="1792x1024", quality="hd", style="vivid", ) return response.data[0].url, response.data[0].revised_prompt ``` Finally, we can get the generated image url from the DALL-E-3 model. Meanwhile, we can also obtain the `revised_prompt` actually used by DALL-E-3 to draw the image, which helps us compare it with our own prompt to see the difference. We also need to implement a `rewrite_query` method. This method will take the list of user historical inputs from Context, then submit it to DeepSeek to rewrite into a complete drawing intention for use when adjusting the drawing prompt or the generated image next time. ```python class ImageGeneration(Workflow): ... @step async def rewrite_query(self, ctx: Context, ev: RewriteQueryEvent) -> None: query_hist_str = await ctx_mgr.get_query_hist(ctx) messages = [ ChatMessage(role="system", content=PROMPT_REWRITE_SYSTEM), ChatMessage(role="user", content=query_hist_str) ] response = await self.deepseek_client.achat(messages) rewritten_prompt = response.message.content await ctx_mgr.set_rewritten_hist(ctx, rewritten_prompt) ``` ### Prompts provided to LLMs The `prompts.py` file contains several system prompts provided to large models: `PROMPT_GENERATE_SYSTEM` is used to prompt the DeepSeek model to generate ink style drawing prompts. This prompt has `{hist_query}` and `{hist_prompt}` placeholders to include user historical requests and the last generated drawing prompt: ```python PROMPT_GENERATE_SYSTEM = """ ## Role You're a visual art designer who's great at writing prompts perfect for DALL-E-3 image generation. ## Task Based on the [image content] I give you, and considering [previous requests], rewrite the prompt to be ideal for DALL-E-3 drawing. ## Length List 4 detailed sentences describing the prompt only - no intros or explanations. ## Context Handling If the message includes [previous prompts], modify them based on the new info. ## Art Style The artwork should be ink-wash style illustrations on slightly yellowed rice paper. ## Previous Requests {hist_query} ## Previous Prompts {hist_prompt} """ ``` `PROMPT_TRANSLATE_SYSTEM` is used to translate the prompts generated in the previous step. The content is relatively simple: ```python PROMPT_TRANSLATE_SYSTEM = """ ## Role You're a professional translator in the AI field, great at turning English prompts into accurate Chinese. ## Task Translate the [original prompt] I give you into Chinese. ## Requirements Only provide the Chinese translation, no intros or explanations. ----------- Original prompt: """ ``` `PROMPT_REWRITE_SYSTEM` is used to rewrite user historical requests. ```python PROMPT_REWRITE_SYSTEM = """ You're a conversation history rewrite assistant. I'll give you a list of requests describing a scene, and you'll rewrite them into one complete sentence. Keep the same description of the scene, and don't add anything not in the original list. """ ``` ### Context operation module Since LlamaIndex Workflow version 1.0 adjusted the api for Context, we need to access `Context.store` to read and save status data. Therefore, I specifically wrote a Context operation module `ctx_manager.py`. The `set_image_prompt` and `get_image_prompt` methods are used to store and retrieve the prompts generated by DeepSeek for drawing. ```python async def set_image_prompt(ctx: Context, image_prompt: str) -> None: await ctx.store.set("image_prompt", image_prompt) async def get_image_prompt(ctx: Context) -> str: image_prompt = await ctx.store.get("image_prompt", "") return image_prompt ``` The `add_query_hist` method will add all user historical requests into a list container in Context. The `get_query_hist` method will take out the historical requests from the container and concatenate them into a string. ```python async def add_query_hist(ctx: Context, user_query: str) -> None: query_hist = await ctx.store.get("query_hist", []) query_hist.append(user_query) await ctx.store.set("query_hist", query_hist) async def get_query_hist(ctx: Context) -> str: query_hist = await ctx.store.get("query_hist", []) query_hist_str = "; ".join(query_hist) return query_hist_str ``` The `set_rewritten_hist` and `get_rewritten_hist` methods are used to store and retrieve the rewritten user drawing intentions. ```python async def set_rewritten_hist(ctx: Context, rewritten_hist: str) -> None: await ctx.store.set("rewritten_hist", rewritten_hist) async def get_rewritten_hist(ctx: Context) -> str: rewritten_prompt = await ctx.store.get("rewritten_hist", "") return rewritten_prompt ``` ### Use Chainlit to make user dialogue interface Chainlit uses lifecycle management to organize code. Among them, the `on_chat_start` method annotated with `@cl.on_chat_start` will be called when the user starts a conversation. The `main` method annotated with `@cl.on_message` is used to respond to the user's single conversation. Since we need to share Context between the user's multiple rounds of dialogues, we need to initialize Workflow and Context in the `on_chat_start` method and store them in `user_session`. This way, they can be reused in the `main` method. ```python @cl.on_chat_start async def on_chat_start(): workflow = ImageGeneration(timeout=300) context = Context(workflow) cl.user_session.set("context", context) cl.user_session.set("workflow", workflow) ``` In the `main` method, besides getting workflow and context, we also need to initialize a `cl.Message` instance. This instance will keep updating content as it gets messages from workflow. Also, if the workflow returns an image generation event, it will be displayed through this Message instance. ```python @cl.on_message async def main(message: cl.Message): workflow = cl.user_session.get("workflow") context = cl.user_session.get("context") msg = cl.Message(content="Generating...") await msg.send() ... ``` We need to display streaming messages returned from both `prompt_generator` and `translate_prompt` nodes in the same Message instance. Therefore, we can use `prompt_result` and `translate_result` to concatenate the current round of messages and update them together with a template. ```python @cl.on_message async def main(message: cl.Message): ... prompt_result = "" translate_result = "" handler = workflow.run(query=message.content, ctx=context) async for event in handler.stream_events(): if isinstance(event, StreamEvent): # # await msg.stream_token(event.delta) match event.target: case "prompt": prompt_result += event.delta case "translate": translate_result += event.delta msg.content = dedent(f""" ### Prompt\n {prompt_result} ### Translate {translate_result} APPROVE? """) await msg.update() ... await handler ``` If the message returned by the workflow is an image drawing message, we can use a `cl.Image` to display the image content. ```python @cl.on_message async def main(message: cl.Message): ... if isinstance(event, StopEvent) and event.target == "image": image = cl.Image(url=event.result["image_url"], name="image1", display="inline") msg.content = f"Revised prompt: \n{event.result['revised_prompt']}" msg.elements = [image] await msg.update() ``` Meanwhile, we can also show the actual `revised_prompt` used by DALL-E-3 to generate the image in the Message, making it convenient for us to compare the accuracy of the image. ### Check the running effect of the workflow Thus far, all project codes have been developed. We can start the Chainlit app through the command line to begin interacting with the workflow: ```bash chainlit run app.py ``` Enter a drawing intention to see the prompt returned by the workflow: ![Enter your drawing idea to see the generated prompt.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-29.png) Enter your drawing idea to see the generated prompt. Image by Author If not satisfied, you can supplement details to make adjustments: ![If you're not happy with it, you can add more details.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-30.png) If you're not happy with it, you can add more details. Image by Author You can see that since we adjusted the prompt used for drawing before the image generation started, this saves a lot of expensive drawing tokens. Of course, this workflow still supports you to continue adjusting the prompt after the image is generated: ![After the image is generated, you can still tweak the prompt.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-31.png) After the image is generated, you can still tweak the prompt. Image by Author --- ## Conclusion The debate about whether to use low-code workflow platforms like dify and n8n to build agent workflows or use coding frameworks to develop agent applications has always existed. Fortunately, the emergence of LlamaIndex Workflow gives us a third option. Its low-level abstraction and simple API make it convenient to develop a production-ready agent workflow, and it can also be customized for various details according to enterprise-level application needs. In today's article, we demonstrated this capability by creating a customized workflow for generating ink painting style images. At the same time, this article itself helps you achieve a beautiful project. Through practicing this project, I have shared several tips for developing agent workflows. These tips condense our experience and insights gained during this period of enterprise-level workflow development. Hopefully, it can help with your multi-agent application development. --- Below is the project source code for this article: [agentic-ai-playground/05\_Image\_Generation\_Workflow at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-7.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-5)](https://github.com/qtalen/agentic-ai-playground/tree/main/05%5FImage%5FGeneration%5FWorkflow?ref=dataleadsfuture.com) --- ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracing Tutorial URL: https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/ Last updated: 2026-04-22T08:25:11.000Z Let's face it - most multi-agent application tutorials online these days are toys. Running them doesn't reliably produce the expected results. So in today's article, I'll walk you through in detail how we use MLflow's latest 3.1 version to trace and monitor agent applications developed based on Qwen 3 models in enterprise-level agent application development workflows. This will give you the ability to develop enterprise-grade high-reliability agent applications. The complete project source code involved in this article is placed at the end of the article for your free reading and modification. --- ## Introduction As a data scientist with years of experience, before the AI era arrived, I had already developed many enterprise-grade algorithmic applications. From my years of experience, measuring whether an algorithmic application is good to use doesn't just depend on whether the application uses the latest technology or has high evaluation metrics. The key is how to ensure that algorithmic applications can stably and reliably provide users with expected results. Namely, what we usually call observability, explainability, and traceability. What do I mean? **Observability:** How does your application run at each step? Are there logs or visual reports to observe? Can developers or administrators monitor the running status at any time? **Explainability:** For each step in the pipeline, can the algorithm explain why a certain input produced this result or caused an error? **Traceability:** If errors occur during code execution or the obtained results deviate too much from expectations, can we accurately locate the cause of the problem and stably reproduce this error to confirm whether the issue is resolved? ![Three key metrics for evaluating the stability of enterprise GenAI applications.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/Monitor_Qwen_3_Agents_with_MLflow-explain.drawio.png) Three key metrics for evaluating the stability of enterprise GenAI applications. Image by Author After entering the AI agent era, with the emergence of numerous multi-agent development frameworks and increasingly higher levels of abstraction, engineers find developing new agent applications convenient. However, tracking and observing the effects during agent runtime becomes more difficult. This leads to situations where we often don't know what the final prompt fed into the large language model is, why we didn't get the desired results, or how multiple agents are orchestrated during runtime. Therefore, we urgently need a tool to help us observe and evaluate agent applications, ensuring we have full control over the entire agent operation process. In the machine learning era, you should have used MLflow to track model training. Fortunately, MLflow recently launched version 3.0, adding tracking and evaluation capabilities for GenAI projects. Moreover, as an open-source project, it can meet data compliance requirements through self-hosted deployment. So in today's article, I will explain in detail how to use MLflow 3.1 to track and monitor my multi-agent applications. ### Why Should You Care? In today's tutorial, I will guide you through the following content: 1. How to install MLflow 3.1 and prepare for agent application tracking. 2. Explain the usage of MLflow 3.1, including annotations, autolog, context manager, and how to handle situations like streaming output. 3. Introduce MLflow's UI interface and basic concepts. 4. How to use MLflow for tracking in Autogen agents and fix bugs in Autogen autolog. 5. Use an Autogen GraphFlow project as an example to demonstrate how to use MLflow in multi-agent projects and record various information needed for tracking. Through today's learning, you will save a lot of technical selection time and be able to proficiently use MLflow 3.x to track and monitor your multi-agents. Let's get started! --- ## Prepare the MLflow Environment ### Install MLflow Server Installing MLflow is relatively simple. In your virtual environment, you can directly use pip to install: ```Bash pip install 'mlflow>=3.1' ``` Since MLflow started focusing on tracking and evaluating GenAI apps from version 3.0, in my experience, version 3.1 has significant changes in API usage compared to version 3.0\. To smoothly check the official website's API documentation and code examples, I recommend installing versions after 3.1. After installation, you can start MLflow with the following command: ```bash mlflow server --host 0.0.0.0 --port 5000 ``` Of course, I recommend starting the MLflow service using docker: ```bash docker pull ghcr.io/mlflow/mlflow:v3.1.1 docker run -d --name mlflow-server \ -p 5000:5000 \ -v $(pwd)/mlruns:/mlflow/mlruns \ ghcr.io/mlflow/mlflow:v3.1.0 \ mlflow ui --host 0.0.0.0 ``` If you are installing in a development environment, simply use the `mlflow ui` command to start the server. At this point, you can access the MLflow UI interface via `http://localhost:5000/`: ![MLflow main interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image.png) MLflow main interface. Image by Author ### Configure Client Code Configuring MLflow's client is relatively straightforward. You can directly add the following code to connect: ```python mlflow.set_tracking_uri("http://localhost:5000") ``` Of course, I recommend configuring via adding `MLFLOW_TRACKING_URI` in environment variables. ```bash MLFLOW_TRACKING_URI=http://localhost:5000/ ``` After configuring both client and server, we can start using MLflow in your openai client code. --- ## Track Your OpenAI Client Code with MLflow ### Use Basic Annotation Method Using MLflow is very simple; you only need one line of code to get started. First, let's start with a basic OpenAI client API call: ```python mlflow.set_experiment("test_openai_tracing") async_client = openai.AsyncOpenAI() async def main(user_query: str) -> str: messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": user_query}, ] response = await async_client.chat.completions.create( model="qwen-plus-latest", temperature=0.7, messages=messages, ) return response.choices[0].message.content ``` Next, we introduce mlflow and add the `mlflow.trace` annotation to the main method. ```python import mlflow @mlflow.trace async def main(user_query: str) -> str: ... ``` I suggest you set up an experiment. If you want to put all project tracking under the same experiment, you can also add a key in the environment variables: ```bash MLFLOW_EXPERIMENT_NAME="test_openai_tracing" ``` Don't forget to start your MLflow server first with `mlflow ui`. Then run the code and open MLflow's UI interface. At this point, we can see the previously tracked record. ![Your first tracking.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-1.png) Your first tracking. Image by Author ### Introduce UI Interface and Some Basic Concepts Next, let's explain some basic concepts combined with the previous tracking: If you open MLflow UI, the first thing displayed is the experiment we are tracking, such as the `test_openai_tracing` we just set in the code. Select the experiment you want to view, click the Traces tab on the top right, and you can view all executed tracking records under the current experiment. You can tag each tracking record in code for easy filtering. Click on the previously executed tracking record and open it, and we can see the detailed information included in this tracking: ![Our Hello World Span, along with its detailed information.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-2.png) Our Hello World Span, along with its detailed information. Image by Author The left side is a tracking event, which MLflow calls a Span. Since we used the annotation method on the main method, the span name here is main. On the right are three tabs: Inputs/Outputs, Attributes, and Events. Since we tracked the main method, Inputs/Outputs show the method’s inputs and outputs respectively. Later, if we track the OpenAI chat API, Inputs will display all parameters passed to the Qwen 3 LLM including messages. Outputs are the messages generated by the large language model. Attributes can record various custom attributes, and what to record is entirely up to you, helping us better document operational information. ![You can use Attributes to help record additional information.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-3.png) You can use Attributes to help record additional information. Image by Author If exceptions are thrown during execution, they will be recorded in the Events tab. If you are calling streaming output, corresponding SSE messages will also be recorded here. ![You can view code execution error messages in the Events interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-4.png) You can view code execution error messages in the Events interface. Image by Author After explaining the UI interface, here is a brief introduction to some basic MLflow concepts: To better organize tracking logs, MLflow's entire tracking system can be viewed as a tree structure. ![MLflow's tracking system can be seen as a tree-like structure.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-5.png) MLflow's tracking system can be seen as a tree-like structure. Image by Author As shown in the previous code, the root nodes of tracking are individual experiments. You can put all tracking of a project into one experiment, or place different iterations of the project into separate experiments. It all depends on the size of your project and the granularity of tracking. Next is trace, representing the tracking of a single code execution. You can consider trace as the entry point of this code execution. A trace contains two independent data structures: `TraceInfo` and `TraceData`. TraceInfo includes duration time, tags, status, and other information to facilitate your filtering. `TraceData` is a collection of Spans. What is Span? After enabling MLflow, you can allow different stages of each code execution to throw events. These events can represent either a method call, the execution of specific code blocks, or exceptions thrown. These events are different Spans. Each Span has a `trace_id`, indicating which trace this Span belongs to. Spans also have parent-child relationships, with child Spans identifying their parent node through `parent_id`. For example, if your main method throws an event, and the `main` method calls the OpenAI `create` method which also throws an LLM event, the Spans corresponding to these two events form a parent-child relationship. Each Span also has its own `SpanType`. For different `SpanTypes`, not only do they display different icons on the UI interface, but some special Spans also look significantly different on the right interface. For instance, LLM SpanType displays the context in a conversational way. Therefore, it is recommended to set different Types for different Spans to better observe program execution. ![Different SpanTypes have distinct icons and interfaces.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-6.png) Different SpanTypes have distinct icons and interfaces. Image by Author ### Use Autolog to Track LLM Calls Earlier, we introduced how to use the `mlflow.trace` annotation to track application code. However, for agent development frameworks, or directly for OpenAI client code, this tracking method is not feasible because we cannot modify the framework source code. And if you add annotations to methods calling client APIs, you cannot finely record what parameters are passed to the API. At this point, we can enable MLflow's autolog feature. Enabling this feature is also very simple; just add one line of code at the beginning of the program. For example, here we want to automatically track calls to the OpenAI API. So we enable OpenAI's autolog: ```python mlflow.openai.autolog() ``` Note that since the principle of the autolog method is to monkey patch the original methods of the corresponding API, you need to ensure that the patched module is imported in advance. For example, you should import OpenAI before enabling OpenAI's autolog. Next, let's demonstrate the autolog effect with a simple OpenAI client call: ```python @mlflow.trace(span_type=SpanType.CHAT_MODEL) async def chatbot(user_query: str, messages: list[dict[str, str]]) -> str: messages.append({ "role": "user", "content": user_query, }) response = await async_client.chat.completions.create( model="qwen-turbo-latest", temperature=0.7, messages=messages, max_tokens=100, ) llm_content = response.choices[0].message.content messages.append({ "role": "assistant", "content": llm_content }) return f"🤖Tony says: {truncate_str(llm_content)}" @mlflow.trace(span_type=SpanType.CHAIN) async def main(): greetings = "Hello, what can I help you with today?" messages = [ {"role": "system", "content": "You are Tony, a fun chatbot."}, {"role": "assistant", "content": greetings}, ] print(f"🤖Tony says: {greetings}") while True: user_query = input(">>> ") if "BYE" in user_query.upper(): break tony_says = await chatbot(user_query, messages) print(tony_says) ``` In this example, we developed a simple chat program using the OpenAI native API. We added the `mlflow.trace` annotation to the `main` and `chatbot` methods. Since this chatbot supports multi-turn conversations, each message sent to the LLM is assembled from historical chat context and the latest user input. Additionally, I truncated the text generated by the LLM in the program. This means that simply tracking the `chatbot` method, you fundamentally don't know what was input to the LLM or what was output. Now, let's add the autolog code and rerun: ```python import mlflow mlflow.set_experiment("test_openai_tracing") mlflow.openai.autolog() ``` Open the UI interface and take a look. You will be pleasantly surprised to find that MLflow not only records the `openai chat completion` API call but also documents the entire conversation message in a dedicated interface: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-7.png) OpenAI Autolog keeps detailed records of the entire conversation. Image by Author ### Track Generator and LLM Streaming Output Besides traditional method calls, in agent applications, we often face situations where we need to record LLM streaming output. In the previous content, I mentioned that through the Events tab on the Span page, we can record SSE events obtained by the OpenAI API. Let's see how it's done. First, write a simple streaming output code: ```python mlflow.openai.autolog() async def predict(query: str) -> AsyncGenerator[tuple[str, Any] | ChatCompletionChunk, None]: messages = [ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": query} ] stream = await async_client.chat.completions.create( model="qwen-plus-latest", temperature=0.8, messages=messages, stream=True ) async for chunk in stream: yield chunk ``` Next, let's go to the Events tab: ![The Events tab records every SSE message.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-8.png) The Events tab records every SSE message. Image by Author You can see that it lists each SSE message received by the client. However, many times, we want to view the concatenated content of all SSE messages in the MLflow interface. Besides the OpenAI client, there are many other methods that also generate content as `generator`s, requiring the `generator` output to be reduced into a complete message. At this point, the `output_reducer` parameter of `mlflow.trace` comes in handy. Before using `output_reducer`, you need to create a reducer method. The method's parameter is the message generated by the `generator`, and the return value is the concatenated text or message: ```python def aggregate_chunks(outputs: list[ChatCompletionChunk]) -> str | None: if not outputs: return None result = "" for chunk in outputs: result += chunk.choices[0].delta.content return result ``` Then, we only need to pass this method through the `output_reducer` parameter in `mlflow.trace`: ```python @mlflow.trace(span_type=SpanType.LLM, output_reducer=aggregate_chunks) async def predict(query: str) -> AsyncGenerator[tuple[str, Any] | ChatCompletionChunk, None]: ... ``` Let's open the MLflow interface and take a look at the predict method's inputs and outputs. We can see that the LLM's streaming messages have been recorded as a fully concatenated text message: ![With the output_reducer parameter, we can stitch streaming outputs into a complete message.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-9.png) With the output\_reducer parameter, we can stitch streaming outputs into a complete message. Image by Author Of course, if you enable OpenAI's autolog, MLflow will automatically concatenate SSE messages from the chat completion API: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-10.png) OpenAI Autolog can also stitch together streaming messages in the Chat interface. Image by Author But for generator methods, output\_reducer is more versatile and allows customizing the concatenated message body. For details, refer to the [official documentation](https://mlflow.org/docs/latest/genai/tracing/app-instrumentation/manual-tracing/fluent-apis?ref=dataleadsfuture.com#streaming). ### Context Manager and Function Calling Finally, we will consider a more complex situation: if you develop an agent program, then the LLM will not only generate responses to user requests but also call specific tools based on user intent and rewrite the execution results before returning them to the user. This process represents multi-step operations within an agent method. We cannot simply use `mlflow.trace` to record method calls, nor can we only use autolog to record OpenAI API calls. At this point, we can use MLflow's context manager to record each step in the agent's running process separately with a Span. Meanwhile, intermediate outputs can be recorded in the Span's attributes, making it easier to track. Next, let's simulate agent execution with a native OpenAI function calling. First, we define a `search_web` method that takes user input as a parameter, uses the Tavily API to search the web, and returns search results: ```python @mlflow.trace(span_type=SpanType.TOOL) async def search_web(query: str) -> str: web_client = AsyncTavilyClient() response = await web_client.search(query) return str(response["results"]) ``` According to the OpenAI API documentation, we also need to convert the tool into a specific structured description: ```python tools = [{ "type": "function", "function": { "name": "search_web", "description": "Find information on the web.", "parameters": { "type": "object", "properties": { "query": { "type": "string", "description": "What you want to search for." } }, "required": ["query"] } } }] _tool_functions = {"search_web": search_web} ``` Then we define a `call_llm` method that calls Qwen 3, passes the message context and callable tools to Qwen, and waits for the model to return the corresponding message, which may contain `tool_calls` or the final result. ```python async def call_llm(messages: list[dict], tools: list[dict] | None = None) \ -> ChatCompletionMessage: response = await async_client.chat.completions.create( model=MODEL_NAME, temperature=0.01, messages=messages, tools=tools, ) return response.choices[0].message ``` Define a `tool_invoke` method. When the message returned by the LLM contains `tool_calls`, we use this method to call the corresponding `tool` to obtain results. ```python async def tool_invoke(message: ChatCompletionMessage, messages: list[dict]) -> list[dict]: result_messages = copy.deepcopy(messages) tool_calls = message.tool_calls for tool_call in tool_calls: function_name = tool_call.function.name tool_func = _tool_functions[function_name] args = json.loads(tool_call.function.arguments) tool_result = await tool_func(**args) result_messages.append({ "role": "tool", "tool_call_id": tool_call.id, "content": tool_result, }) return result_messages ``` Finally, the `search_agent` method acts as the agent, containing the call order of the preceding methods. We first implement the basic code logic without MLflow tracking. ```python @mlflow.trace(span_type=SpanType.AGENT) async def search_agent(query: str) -> str: messages = [{ "role": "system", "content": "You are a helpful assistant, and you use search_web tool to find information on the web.", }, { "role": "user", "content": query, }] message = await call_llm(messages, tools) if len(message.content) > 0: return message.content messages.append(message.model_dump()) messages = await tool_invoke(message, messages) message = await call_llm(messages) return message.content ``` Executing the `search_agent` method shows that the agent searches the web and returns organized results based on our provided tasks: ![Autolog's tracking records. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-11.png) Autolog's tracking records. Image by Author We summarize the several steps of executing the `search_agent` code: deciding which tool to call based on user requests, calling the specific tool, and generating corresponding output based on the tool's return. ![A typical Agent invocation process.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-12.png) A typical Agent invocation process. Image by Author Next, we use MLflow's context manager to track these three steps separately and additionally record the input and output data of each phase. First, we add two tags to the current trace: the date of execution and the model name used during execution. This will facilitate subsequent trace filtering and quickly locate the desired records: ```python mlflow.update_current_trace( tags={ "date": date.today().strftime("%Y%M%d"), "model": MODEL_NAME } ) ``` Next, start defining the first Span, named `get_tool_calls`. Simultaneously, we record the user request and the model's returned `message.content` as Inputs/Outputs: ```python with mlflow.start_span(name="get_tool_calls", span_type=SpanType.LLM) as span: span.set_inputs({ "query": query }) messages = [{ "role": "system", "content": "You are a helpful assistant, and you use search_web tool to find information on the web.", }, { "role": "user", "content": query, }] message = await call_llm(messages, tools) if len(message.content) > 0: span.set_outputs({ "results": message.content }) return message.content messages.append(message.model_dump()) span.set_outputs({ "tool_calls": message.tool_calls, }) span.set_attributes({ "num_of_tool_calls": len(message.tool_calls), }) ``` Define another Span, recording the `tool_calls` and the results obtained from calling the tool as the Span's Inputs/Outputs: ```python with mlflow.start_span(name="invoke_tools", span_type=SpanType.TOOL) as span: span.set_inputs({ "tool_calls": message.tool_calls }) messages = await tool_invoke(message, messages) tool_call_results = messages[-1: -1 - len(message.tool_calls)] span.set_outputs({ "tool_call_results": tool_call_results }) span.set_attributes({ "num_of_tool_call_results": len(tool_call_results), }) ``` Finally, define a Span named `reflect_tool_calls`, recording the final copy generated by the large language model based on the tool's return results: ```python with mlflow.start_span(name="reflect_tool_calls", span_type=SpanType.LLM) as span: span.set_inputs({ "messages": messages, }) message = await call_llm(messages) span.set_outputs({ "answer": message.content }) ``` We check the effect of custom Spans through the MLflow interface. We can see that the parent Span and three custom Spans have been recorded. By clicking on each Span, we can see the complete execution process of the agent on the right, giving us a clear understanding, right? ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-13.png) By customizing Spans, you can see the intermediate steps of the Agent's execution. Image by Author So far, through a few simple OpenAI client practices, we have basically mastered the usage of MLflow. However, in enterprise-level agent applications, we can't start from basic OpenAI code but use higher-abstraction agent frameworks to complete application development. So next, I will use Autogen's `GraphFlow` workflow application as an example to show you how to use MLflow for tracking and observing agent code in enterprise-level application scenarios. --- ## Enhance Observability and Explainability of Autogen Agents with MLflow Currently, my team is using Autogen to build enterprise-level agent applications. If you want to know how this happened, you can read my article: [Build AutoGen Agents with Qwen3: Structured Output & Thinking ModeSave yourself 40 hours of trial and error![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-33.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Autogen_Qwen3_cover_2-1.webp)](https://www.dataleadsfuture.com/build-autogen-agents-with-qwen3-structured-output-thinking-mode/) In the following content, I will start from tracking a simple `AssistantAgent` and proceed to the practice of Autogen `GraphFlow`, showing you how we perform effect tracking in agent applications. ### Fix MLflow 3.1 Bug in Autogen Autolog Before we begin, it needs to be noted that Autogen officially provides a solution for tracking agent applications. But if you follow the official tutorial to deploy the OpenTelemetry service, write the code, and execute it, you will get such an interface. ![Messages generated using the OpenTelemetry solution recommended by Autogen.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-14.png) Messages generated using the OpenTelemetry solution recommended by Autogen. Image by Author Clicking inside, you will see many Span-like structures but no useful information: ![Apart from the agent's name, this message contains no useful information. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-15.png) Apart from the agent's name, this message contains no useful information. Image by Author This is why we use MLflow instead of the official recommended solution today: we need an organized and clearly informative tracking tool. In fact, MLflow also provides autolog for the Autogen framework. You only need to introduce one line of code to start tracking agent execution: ```python mlflow.autogen.autolog() ``` Unfortunately, as of writing this article, using autolog to record agents with function calling in Autogen 0.6.1 version will result in the following error: ```bash WARNING mlflow.utils.autologging_utils: Encountered unexpected error during autogen autologging: 2 validation errors for ChatMessage content.str Input should be a valid string [type=string_type, input_value={'content': [{'content': ...ExecutionResultMessage'}, input_type=dict] For further information visit https://errors.pydantic.dev/2.11/v/string_type content.list[tagged-union[TextContentPart,ImageContentPart,AudioContentPart]] Input should be a valid list [type=list_type, input_value={'content': [{'content': ...ExecutionResultMessage'}, input_type=dict] For further information visit https://errors.pydantic.dev/2.11/v/list_type ``` The reason for this error is that when MLflow monkey patches Autogen’s API and finds that the message type returned by the LLM call is `FunctionExecutionResultMessage`, it calls the `message.model_dump()` method to write into the `ChatMessage`'s `content` attribute. ![When the message is a FunctionExecutionResultMessage, MLflow will call model_dump to write the message. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-16.png) When the message is a FunctionExecutionResultMessage, MLflow will call model\_dump to write the message. Image by Author However, the `content` attribute only accepts values of type `str` and `list`, while the `model_dump()` method returns a `dict` type, thus causing validation errors. Therefore, before continuing to track Autogen applications, we need to fix this bug. Here is the tracking record after the fix: ![After fixing the bug, use MLflow to track the Autogen application interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-17.png) After fixing the bug, use MLflow to track the Autogen application interface. Image by Author As you can see, MLflow has well-recorded the execution steps of Autogen AssistantAgent, including detailed LLM inputs, outputs, and various parameters. Let me explain how I fixed this bug. Previously, we discussed that the reason for the Pydantic validation error was that the `content` attribute of `ChatMessage` did not accept `dict` type values. Therefore, we need to modify `ChatMessage` to support `dict` types for the `content` attribute. Since `ChatMessage` is referenced by the `mlflow.autogen.chat` module, modifying it using subclasses is not feasible. A more viable approach is to monkey patch `ChatMessage` and then replace the reference in autogen autolog. I will place the monkey patching code for `ChatMessage` in the `autogen_patching.py` file. Since the original `ChatMessage` code and its calling code are dynamically loaded in the autolog method, before modifying the `ChatMessage` code, we need to load autolog first: ```python import mlflow TARGET_MODULE = "mlflow.types.chat" ORIGINAL_CLASS_NAME = "ChatMessage" BASE_CLASS = "BaseModel" mlflow.autogen.autolog() ``` Then, we re-implement the `ChatMessage` class and add `dict` type to the `content` attribute: ```python module = importlib.import_module(TARGET_MODULE) BaseModel = getattr(module, ORIGINAL_CLASS_NAME) class ChatMessage(BaseModel): role: str content: str | list | dict | None = None ``` With the help of DeepSeek R1, I wrote a utility for monkey patching ☺️: ```python class ClassReplacer: def __init__( self, target_module: str = TARGET_MODULE, original_class_name: str = ORIGINAL_CLASS_NAME, new_class: Type = None, ): self._target_module = target_module self._original_class_name = original_class_name self._new_class = new_class self._module = importlib.import_module(self._target_module) self._original_class = getattr(self._module, original_class_name) def apply(self): for mod_name, mod in list(sys.modules.items()): if mod is None or not isinstance(mod, ModuleType): continue if hasattr(mod, self._original_class_name): current_ref = getattr(mod, self._original_class_name) if current_ref is self._original_class: setattr(mod, self._original_class_name, self._new_class) ``` In this tool, we first replace the original Class with the re-implemented Class, and then replace all references. Introduce the written `autogen_patching`, and then run our Autogen agent code. Bingo, it runs normally now. ### Project Practice: Tracking Autogen GraphFlow Application Where MLflow truly shines is in AI applications based on multi-agent frameworks. Due to the high level of abstraction and flexible orchestration among agents in these applications, it is challenging to understand the execution order of various agents during actual code execution or what the final prompts fed into the LLM are in each step. However, with the help of MLflow, all this becomes manageable. In the following simple project practice, I will directly use Autogen's official GraphFlow project example to demonstrate how to track and observe in multi-agent applications. Additionally, with the assistance of MLflow, we will fix a hard-to-spot bug in the official example. This is an idea generation workflow project containing generator, reviewer, and summary nodes. The generator generates a series of ideas based on user requests, the reviewer reviews the feasibility of these ideas and decides whether the generator needs to regenerate them, and the summary ultimately compiles the final summary. The flowchart of the entire project is as follows: ![GraphFlow project's flowchart.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-18.png) GraphFlow project's flowchart. Image by Author Then, we create three agents respectively. Generator agent: ```python generator = AssistantAgent( "generator", model_client=model_client, system_message=""" Generate a list of creative ideas. """ ) ``` Reviewer agent: ```python reviewer = AssistantAgent( "reviewer", model_client=model_client, system_message=""" Review ideas and provide feedbacks, or just 'APPROVE' for final approval. """ ) ``` Summary agent: ```python summarizer_core = AssistantAgent( "summary", model_client=model_client, system_message=""" Summarize the user request and the final feedback. """ ) ``` I previously wrote an article discussing the issue of position bias caused by passing too many messages to the LLM, leading to poor generation quality due to the inability to understand user intent: [Fixing the Agent Handoff Problem in LlamaIndex’s AgentWorkflow SystemThe position bias in LLMs is the root cause of the problem![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-34.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/p60Imf7mlg8tNL8kC9VdJSw8AsixrM-2.webp)](https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/) Autogen introduces a wrapper called `MessageFilterAgent`, which filters messages input to the wrapped agent, thereby avoiding position bias caused by inputting too many messages. Take the following code as an example. Since the generator and reviewer agents engage in multiple rounds of discussion loops to reach a final conclusion, we only need to send the last message to the summary. Whether to filter the last message from the reviewer will be discussed later: ```python filtered_summarizer = MessageFilterAgent( name="summary", wrapped_agent=summarizer_core, filter=MessageFilterConfig( per_source=[ PerSourceFilter(source="user", position="first", count=1), PerSourceFilter(source="reviewer", position="last", count=1), ] ) ) ``` Finally, we use GraphFlow to integrate the above agent nodes into a DAG workflow structure: ```python builder = DiGraphBuilder() builder.add_node(generator).add_node(reviewer).add_node(filtered_summarizer) builder.add_edge(generator, reviewer) builder.add_edge(reviewer, filtered_summarizer, condition=lambda msg: "APPROVE" in msg.to_model_text()) builder.add_edge(reviewer, generator, condition=lambda msg: "APPROVE" not in msg.to_model_text()) builder.set_entry_point(generator) graph = builder.build() flow = GraphFlow( participants=builder.get_participants(), graph=graph, ) ``` Check the effect of the agent’s operation: ![The execution results of the Autogen GraphFlow sample code on the official website did not yield any meaningful insights.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-19.png) The execution results of the Autogen GraphFlow sample code on the official website did not yield any meaningful insights. Image by Author It seems the workflow ran normally, but it also appears it didn't. The summary agent outputs the final summary, yet it doesn’t seem to summarize any substantial content: Isn't this the norm in our daily development of agent applications? But it indeed doesn't meet the usability requirements of enterprise-grade GenAI applications. Let's examine where the problem lies. Remember the modified autogen autolog script from earlier? Let's introduce it and rerun the entire application. ```python import utils.autogen_patching ``` This time, let's observe what message the summary agent finally passed to the Qwen 3 model via the MLflow interface: ![Only "APPROVE" is passed to the summary agent, not the content generated by the generator.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-20.png) Only "APPROVE" is passed to the summary agent, not the content generated by the generator. Image by Author Did you spot the issue? We only passed the user's request and the APPROVE message generated by the reviewer agent to the large language model, without including the creative content generated by the generator agent. This is because in the configuration of `MessageFilterAgent`, we indeed configured it to retain the first message from the user and the last message from the reviewer: ```python filtered_summarizer = MessageFilterAgent( name="summary", wrapped_agent=summarizer_core, filter=MessageFilterConfig( per_source=[ PerSourceFilter(source="user", position="first", count=1), PerSourceFilter(source="reviewer", position="last", count=1), ] ) ) ``` But what we really need is the last message generated by the generator, so we need to change the code to retain the last message generated by the generator: ```python filtered_summarizer = MessageFilterAgent( name="summary", wrapped_agent=summarizer_core, filter=MessageFilterConfig( per_source=[ PerSourceFilter(source="user", position="first", count=1), PerSourceFilter(source="generator", position="last", count=1), ] ) ) ``` Then run the code again, and observe the tracking interface: ![The generator's information was passed to the summary agent, ultimately producing a comprehensive summary.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-21.png) The generator's information was passed to the summary agent, ultimately producing a comprehensive summary. Image by Author As you can see, this time the summary got the user's request and the final approved message from the generator, and generated a content-rich summary. How about that? With the help of MLflow, problems that were originally hard to detect can now be clearly observed through the interface and resolved. MLflow once again proves its value! Are we done? Hold on a bit longer. Besides recording the input and output messages of large language models, to conveniently track logs of multiple agent executions and document performance, we need to record additional information. We still use the context manager method to wrap the application execution code in the main method. As usual, we first add date and model tags to the trace, making it easier to filter traces you focus on later: ```python mlflow.update_current_trace( tags={ "date": date.today().strftime('%Y%m%d'), "model": model_client_config.get('model'), } ) ``` As a measure of agent-generated performance, in this project, we are concerned with the number of cycles between the generator and reviewer. Fewer cycles indicate lower token usage costs, so we can record how many times the generator created and the corresponding large language model settings in the root Span: ```python generates = [msg for msg in result.messages if msg.source == "generator"] final_generate = generates[-1] summary = [msg for msg in result.messages if msg.source == "summary"][0] current_span.set_outputs({ "generate": final_generate.content, "summary": summary.content, }) current_span.set_attributes({ "rounds": str(len(generates)), **model_client_config, }) ``` OK, now open the MLflow interface again, and you can see the customized data just recorded. Of course, you can also add some tracking data according to your needs. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/07/image-22.png) You can view the additional information we recorded through the Attributes interface. Image by Author Thus far, I have explained how to conduct effect tracking and observation in Autogen multi-agent applications. Isn't it simple? --- ## Conclusion Today, we can learn multi-agent application development methods through various tutorials, but rarely do people discuss tracking and observation solutions for agent applications. For enterprise-level intelligent applications, as the abstraction levels of various agent development frameworks become higher, using tools to enhance the observability, explainability, and traceability of applications is becoming increasingly important. In today's article, I explained how to use MLflow 3.x version to track and observe agent applications, covering the use of annotations, autolog, and context managers, as well as a detailed tutorial on how to track Autogen GraphFlow. In upcoming articles, I will continue to explain how to manage the effectiveness of different versions of prompts and how to evaluate the generation effects of agents. Stay tuned. Thank you for subscribing. I hope my Agentic AI series tutorials make you feel they are worth more than their price. Feel free to leave comments for discussion, and I will reply as soon as possible. Here is the source code for this article: [agentic-ai-playground/04\_Monitoring\_Qwen3\_Agents\_with\_MLflow\_3 at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-6.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-4)](https://github.com/qtalen/agentic-ai-playground/tree/main/04%5FMonitoring%5FQwen3%5FAgents%5Fwith%5FMLflow%5F3?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Build AutoGen Agents with Qwen3: Structured Output & Thinking Mode URL: https://www.dataleadsfuture.com/build-autogen-agents-with-qwen3-structured-output-thinking-mode/ Last updated: 2026-07-27T01:41:56.000Z *Disclaimer: This post contains affiliate links to Coursera. If you click and enroll, I may earn a small commission at no extra cost to you. Thank you for your support!* This article will walk you through integrating AutoGen with Qwen3, including how to enable structured output for Qwen3 in AutoGen and manage Qwen3's thinking mode capabilities. If you're in a hurry for the solution, you can skip the "how-to" sections and jump straight to the end, where I've shared all the source code. Feel free to use and modify it without asking for permission. Autogen has stopped updating, so I’ve also prepared a Microsoft Agent Framework version of the solution for you. Click here to learn more: [Make Microsoft Agent Framework’s Structured Output Work With Qwen and DeepSeek ModelsThings You Always Have to Do When Switching a Framework![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-51.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Agent_Framework_cover.webp)](https://www.dataleadsfuture.com/make-microsoft-agent-frameworks-structured-output-work-with-qwen-and-deepseek-models/) --- ## Introduction As enterprises begin deploying Qwen3 models, corresponding agent frameworks must adapt to fully utilize Qwen3's capabilities. For the past two months, my team and I have been working on a large-scale project using AutoGen. Like LlamaIndex Workflow, this event-driven agent framework allows our agents to integrate better with enterprise message pipelines, leveraging the full power of our data processing architecture. If you're also interested in LlamaIndex Workflow, I've written two articles about it: [Deep Dive into LlamaIndex Workflow: Event-driven LLM architectureWhat I think about the progress and shortcomings after practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-28.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LLamaIndex_Workflow-3.webp)](https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/) [Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration SolutionAnd fix the issue where the agent can’t continue with past requests![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-29.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4-1-4.webp)](https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/) We've spent considerable time with Qwen3, experimenting with various approaches and even consulting directly with the Qwen team on certain options. I'm confident this article will help you, even if you're not using Qwen series models or AutoGen specifically. The problem-solving approaches are universal, so you'll save significant time. ### Why should I care? There's an old Chinese saying: "A craftsman must first sharpen his tools." To fully benefit from the latest technology's performance improvements and development conveniences, integrating models into existing systems is the first step. This article will cover: - Creating an OpenAI-like client that lets AutoGen connect to Qwen3 via OpenAI API. - Exploring AutoGen's structured output implementation and alternative approaches, ultimately adding structured output support for Qwen3. - Supporting Qwen3's `extra_body` parameters in AutoGen by controlling the thinking mode toggle. - Finally, we'll put these lessons into practice with an article-summarizing agent project. Let's begin! --- ## Step 1: Building an OpenAILike Client ### Testing the official OpenAI client We know both Qwen and DeepSeek models support OpenAI API calls. AutoGen provides a `OpenAIChatCompletionClient` class for GPT series models. Can we use it to connect to public cloud or privately deployed Qwen3 models? Unfortunately not. When we tried: ```python original_model_client = OpenAIChatCompletionClient( model="qwen-plus-latest", base_url=os.getenv("OPENAI_BASE_URL") ) agent = AssistantAgent( name="assistant", model_client=original_model_client, system_message="You are a helpful assistant." ) ``` We encountered an error: ![When using OpenAIChatCompletionClient, you need to specify an OpenAI model.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-2.png) When using OpenAIChatCompletionClient, you need to specify an OpenAI model. Image by Author Checking the [\_model\_client.py](https://microsoft.github.io/autogen/dev//reference/python/autogen%5Fcore.models.html?ref=dataleadsfuture.com#autogen%5Fcore.models.ModelInfo) file reveals that `OpenAIChatCompletionClient` only supports GPT series models, plus some Gemini and Claude models - no promises for others. But this isn't new. Remember when the OpenAI client last restricted model types? Exactly - LlamaIndex's OpenAI client had similar limitations, but the community provided an OpenAILike client as a workaround. [How to Connect LlamaIndex with Private LLM API DeploymentsWhen your enterprise doesn’t use public models like OpenAI![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-30.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LlamaIndex_LLM_API.drawio-2.png)](https://www.dataleadsfuture.com/how-to-connect-llamaindex-with-private-llm-api-deployments/) Our solution here is similar: we'll build an OpenAILike client supporting Qwen (and DeepSeek) series models. ### Trying the model\_info parameter Checking the [API docs](https://microsoft.github.io/autogen/dev//reference/python/autogen%5Fext.models.openai.html?ref=dataleadsfuture.com) reveals a `mode_info` parameter: *"Required if the model name is not a valid OpenAI model."* So `OpenAIChatCompletionClient` can support non-OpenAI models if we provide the model's own `mode_info`. For Qwen3 models on public cloud, `qwen-plus-lastest` and `qwen-turbo-latest` are the newest. I'll demonstrate with qwen-plus-latest: ```python original_model_client = OpenAIChatCompletionClient( model="qwen-plus-latest", base_url=os.getenv("OPENAI_BASE_URL"), model_info={ "vision": False, "function_calling": True, "json_output": True, "family": 'qwen', "structured_output": True, "multiple_system_messages": False, } ) ... async def main(): await Console( agent.run_stream(task="Hi, could you introduce yourself? Are you Qwen3?") ) ``` ![After adding custom model_info, the client successfully connected to the Qwen3 model. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-3.png) After adding custom model\_info, the client successfully connected to the Qwen3 model. Image by Author The model connects successfully and generates content normally. But this creates new headaches - I don't want to copy-paste `model_info` constantly, nor do I care about its various options. Solution? Time to implement our own OpenAILike that encapsulates this information. ### Implementing OpenAILikeChatCompletionClient We'll implement this through inheritance. The code lives in `utils/openai_like.py`. By inheriting from `OpenAIChatCompletionClient`, we'll automate `model_info` handling. First, we compile all potential models' `model_info` into a dict, with a default `model_info` for unlisted models. ```python _MODEL_INFO: dict[str, dict] = { ... "qwen-plus-latest": { "vision": False, "function_calling": True, "json_output": True, "family": ModelFamily.QWEN, "structured_output": True, "context_window": 128_000, "multiple_system_messages": False, }, ... } DEFAULT_MODEL_INFO = { "vision": False, "function_calling": True, "json_output": True, "family": ModelFamily.QWEN, "structured_output": True, "context_window": 32_000, "multiple_system_messages": False, } ``` In `__init __`, we check if users provided `model_info`. If not, we look up the model parameter in our config, falling back to default if missing. Since `OpenAIChatCompletionClient` requires users to provide `base_url`, we've optimized this too: if missing, we'll pull from `OPENAI_BASE_URL` or `OPENAI_API_BASE` environment variables. Our final `__init __` method looks like: ```python class OpenAILikeChatCompletionClient(OpenAIChatCompletionClient): def __init__(self, **kwargs): self.model = kwargs.get("model", "qwen-max") if "model_info" not in kwargs: kwargs["model_info"] = _MODEL_INFO.get(self.model, DEFAULT_MODEL_INFO) if "base_url" not in kwargs: kwargs["base_url"] = os.getenv("OPENAI_BASE_URL") or os.getenv("OPENAI_API_BASE") super().__init__(**kwargs) ``` Let's test our `OpenAILikeChatCompletionClient`: ```python model_client = OpenAILikeChatCompletionClient( model="qwen-plus-latest" ) agent = AssistantAgent( name="assistant", model_client=model_client, system_message="You are a helpful assistant." ) ``` Perfect! Just specify the model and we're ready to use the latest Qwen3. --- If you want to skip the usual trial and error and jump straight into an enterprise AI career as fast as possible, I highly recommend checking out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/Agamy1?ref=dataleadsfuture.com). It gives you way more structured guidance to get there. --- ## Step 2: Supporting structured\_output Structured\_output specifies a pydantic BaseModel-derived class as standard output. This provides consistent, predictable output formats for more precise agent messaging. For enterprise applications using frameworks and models, structured\_output is essential. ### AutoGen's structured\_output implementation AutoGen supports structured\_output - just implement a pydantic `BaseModel` class and pass it via `output_content_type` to AssistantAgent. The agent's response then becomes a `StructuredMessage` containing structured output. Let's test Qwen3's structured\_output capability. Following official examples, we'll create a sentiment analysis agent. First, define a data class: ```python class AgentResponse(BaseModel): thoughts: str response: Literal["happy", "sad", "neutral"] ``` Then pass this class to the agent via `output_content_type`: ```python structured_output_agent = AssistantAgent( name="structured_output_agent", model_client=model_client, system_message="Categorize the input as happy, sad, or neutral following json format.", output_content_type=AgentResponse ) ``` Running this agent produces an error because the model's JSON output doesn't match our class definition, suggesting the model didn't receive our parameters: ![The model's JSON output doesn't match our class definition.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-4.png) The model's JSON output doesn't match our class definition. Image by Author Why? Does Qwen not support structured\_output? To answer, we need to understand AutoGen's structured\_output implementation. When working directly with LLMs, structured\_output typically adjusts the chat completion API's `response_format` parameter. Having modified `OpenAIChatCompletionClient` earlier, we check its code for structured\_output references and find this comment: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-5.png) Comment of structured\_output in OpenAIChatCompletionClient. Image by Author This suggests that with structured\_output, OpenAI client's `response_format` parameter is set to: ```python { "type": "json_schema", "json_schema": { "name": "name of the schema, must be an identifier.", "description": "description for the model.", "schema": "", "strict": False, # or True }, } ``` But Qwen3's documentation shows its response\_format only supports `{"type": "text"}` and `{"type": "json_object"}`, not `{"type": "json_schema"}`. Does this mean Qwen3 can't do structured\_output? Not necessarily. The essence of structured\_output is getting the model to output JSON matching our schema. Without `response_format`, we have other solutions. ### Implementing structured\_output via function calling Returning to Python's nature: in Python, all classes are callable objects like functions, including pydantic `BaseModel` classes. Can we leverage this with LLM function calling for structured\_output? Absolutely - by treating data classes as special functions. Let's modify our agent. Instead of `output_content_type`, we'll use tools parameter, passing `AgentResponse` as a tool. The model will then call this tool for output: ```python function_calling_agent = AssistantAgent( name="function_calling_agent", model_client=model_client, system_message="Categorize the input as happy, sad, or neutral following json format.", tools=[AgentResponse], ) ``` Results: ![The model outputs JSON perfectly when using function calling.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-6.png) The model outputs JSON perfectly when using function calling. Image by Author The model outputs JSON matching our class's schema perfectly. This works! However, with multiple tools, the agent sometimes ignores the data class tool, outputting freely. Is there a more stable approach? Let's think deeper: structured\_output's essence is getting JSON matching our schema. Can we leverage that directly? ### Making the model output according to json\_schema Having the model output according to `json_schema` is entirely feasible. AutoGen previously used response\_format's `json_schema`, but we can also specify the schema directly in `system_prompt`: ```python json_schema_agent = AssistantAgent( name="json_schema_agent", model_client=model_client, system_message=dedent(f""" Categorize the input as happy, sad, or neutral, And follow the JSON format defined by the following JSON schema: {AgentResponse.model_json_schema()} """) ) ``` Results: ![When we directly specify json_schema in the system_prompt, the model outputs perfect JSON content.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-7.png) When we directly specify json\_schema in the system\_prompt, the model outputs perfect JSON content. Image by Author The agent outputs JSON matching our schema. Understanding the principles makes structured\_output implementation straightforward. We can further convert JSON output back to our data class for code processing: ```python result = await Console(json_schema_agent.run_stream(task="I'm happy.")) structured_result = AgentResponse.model_validate_json( result.messages[-1].content ) print(structured_result.thoughts) print(structured_result.response) ``` ![We can manually convert JSON text into Pydantic data classes.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-8.png) We can manually convert JSON text into Pydantic data classes. Image by Author Perfect. But specifying `json_schema` in `system_prompt` is cumbersome. Can we make Qwen3 agents support `output_content_type` directly? ### Making Qwen3 support AutoGen's output\_content\_type parameter Our ultimate goal is framework-level Qwen3 support for structured\_output via `output_content_type`, without changing AutoGen usage. Earlier we saw `output_content_type` works via `response_format`, set when agents call OpenAI Client's `create` or `create_stream` methods. We also know modifying `system_prompt` to explicitly specify `json_schema` produces stable structured output. Having implemented `OpenAILikeChatCompletionClient`, can we override `create` and `create_stream` to modify `system_prompt`? Let's do it. First, `add _append_json_schema` to `OpenAILikeChatCompletionClient`. This finds the first message in the sequence and appends `json_schema` instructions to `system_prompt`: ```python class OpenAILikeChatCompletionClient(OpenAIChatCompletionClient): ... def _append_json_schema(self, messages: Sequence[LLMMessage], json_output: BaseModel) -> Sequence[LLMMessage]: messages = copy.deepcopy(messages) first_message = messages[0] if isinstance(first_message, SystemMessage): first_message.content += dedent(f"""\ Your output must adhere to the following JSON schema format, without any Markdown syntax, and without any preface or explanation: {json_output.model_json_schema()} """) return messages ``` Then override `create` and `create_stream` to call `_append_json_schema` first, while clearing `json_output` to prevent AutoGen from setting `response_format`: ```python class OpenAILikeChatCompletionClient(OpenAIChatCompletionClient): ... @override async def create( self, messages: Sequence[LLMMessage], *, tools: Sequence[Tool | ToolSchema] = [], json_output: Optional[bool | type[BaseModel]] = None, extra_create_args: Mapping[str, Any] = {}, cancellation_token: Optional[CancellationToken] = None, ) -> CreateResult: if json_output is not None and issubclass(json_output, BaseModel): messages = self._append_json_schema(messages, json_output) json_output = None result = await super().create( messages=messages, tools=tools, json_output=json_output, extra_create_args=extra_create_args, cancellation_token=cancellation_token ) return result ``` Our `OpenAILikeChatCompletionClient` modifications are complete. Since we modified underlying methods, users' `output_content_type` usage remains unchanged. Let's test with a new agent, setting neither tools nor requiring `system_prompt` modifications - just `output_content_type` as per AutoGen docs: ```python structured_output_agent = AssistantAgent( name="structured_output_agent", model_client=model_client, system_message="Categorize the input as happy, sad, or neutral following json format.", output_content_type=AgentResponse ) ``` Now the agent outputs correct JSON directly: ![AutoGen generated StructuredMessage this time.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-9.png) AutoGen generated StructuredMessage this time. Image by Author Notice the message type is `StructuredMessage` \- the agent directly produced our data class, confirmed by isinstance checks: ```python result = await Console( structured_output_agent.run_stream(task="I'm happy") ) print(isinstance(result.messages[-1].content, AgentResponse)) ``` With these changes, AutoGen can correctly generate structured messages in multi-agent systems using Qwen3\. These modifications also work for older Qwen models and DeepSeek series. --- After building your Autogen agent application, you'll want more than just functionality - you'll want robustness, observability, and traceability. With MLflow 3.1, we're happy to say we've achieved this. Check out the article below to learn more: [Monitoring Qwen 3 Agents with MLflow 3.x: End-to-End Tracking TutorialEnhance your multi-agent application’s observability, explainability and Traceability![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-35.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Monitoring-Qwen-3-Agents-with-MLflow-3.x.webp)](https://www.dataleadsfuture.com/monitoring-qwen-3-agents-with-mlflow-3-x-end-to-end-tracking-tutorial/) --- ## Step 3: Supporting Thinking Mode ### Parameters for enabling/disabling thinking mode In previous examples, we used public cloud Qwen3 models (qwen-plus-latest), intentionally ignoring Qwen3's new thinking capability. Enterprise applications typically use privately deployed `qwen3-235b-a22b` or `qwen3-30b-a3b` models. These open-source models differ by defaulting to thinking mode. Before answering, the model performs Chain of Thoughts (CoT) reasoning, significantly improving performance - similar to DeepSeek-R1 or QwQ models. But in multi-model applications, we sometimes want to disable thinking to reduce token usage and latency. Qwen3's documentation shows the `extra_body={"enable_thinking": xxx}` parameter controls thinking mode. Let's test this during client creation: ```python model_client = OpenAILikeChatCompletionClient( model="qwen3-30b-a3b", extra_body={"enable_thinking": False} ) ... async def main(): await Console( agent.run_stream(task="I have nothing but money.") ) ``` Surprisingly, this parameter has no effect: ![Adding the extra_body parameter directly won't have any effect.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-10.png) Adding the extra\_body parameter directly won't have any effect. Image by Author Why? Again, we examine AutoGen's source code. ### How AutoGen handles parameters In `OpenAIChatCompletionClient`, `__init__` calls `_create_args_from_config` to initialize parameters stored in `self._create_args`. Then `_process_create_args` merges `self._create_args` with create's parameters before sending to the model. ![How AutoGen handles parameters.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/autogen_with_qwen3.drawio.png) How AutoGen handles parameters. Image by Author So we can override `_process_create_args` in `OpenAILikeChatCompletionClient` to see what parameters AutoGen sends to Qwen3\. Here we'll just examine `self._create_args`: ```python class OpenAILikeChatCompletionClient(OpenAIChatCompletionClient): ... def _process_create_args( self, messages: Sequence[LLMMessage], tools: Sequence[Tool | ToolSchema], json_output: Optional[bool | type[BaseModel]], extra_create_args: Mapping[str, Any], ) -> CreateParams: print(self._create_args) params = super()._process_create_args( messages=messages, tools=tools, json_output=json_output, extra_create_args=extra_create_args ) return params ``` For comparison, we'll add a `temperature` parameter (which GPT models support): ```python model_client = OpenAILikeChatCompletionClient( model="qwen3-30b-a3b", temperature=0.01, extra_body={"enable_thinking": False} ) ``` The output shows Qwen3 receives `model` and `temperature` parameters, but not `extra_body`. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-11.png) The output shows Qwen3 receives `model` and `temperature` parameters, but not `extra_body`. Image by Author `_create_args_from_config` checks if parameters are GPT-supported, ignoring others. Since `extra_body` is Qwen3-specific, it's ignored. ### Adding extra\_body support Don't worry - we'll add `extra_body` to `self._create_args`. First, let's confirm thinking mode disabled: ![The LLM no longer goes through the thinking process before generating the final result.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-12.png) The LLM no longer goes through the thinking process before generating the final result. Image by Author Supporting `extra_body` is simple. Knowing how `self._create_args` is generated, we just add `extra_body` after parent class initialization. `_create_args_from_config` handles constructor parameters, but as a standalone function, it's not easily overridden. Instead, in `OpenAILikeChatCompletionClient.__init__`, we'll add our parameters after `self._create_args` is created: ```python class OpenAILikeChatCompletionClient(OpenAIChatCompletionClient): def __init__(self, **kwargs): self.model = kwargs.get("model", "qwen-max") if "model_info" not in kwargs: kwargs["model_info"] = _MODEL_INFO.get(self.model, DEFAULT_MODEL_INFO) if "base_url" not in kwargs: kwargs["base_url"] = os.getenv("OPENAI_BASE_URL") or os.getenv("OPENAI_API_BASE") super().__init__(**kwargs) for key in extra_kwargs: # Add the model-specific extension parameters for Qwen3 in self._create_args if key in kwargs: self._create_args[key] = kwargs[key] ``` Now let's test by adding `extra_body={"enable_thinking": False}` during `model_client` creation: ```python model_client = OpenAILikeChatCompletionClient( model="qwen3-30b-a3b", temperature=0.01, extra_body={"enable_thinking": False} ) ``` Checking `self._create_args` again, `extra_body` is successfully included: ![extra_body is successfully included.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-13.png) `extra_body` is successfully included. Image by Author ### Alternative thinking mode control methods Beyond `extra_body`, official documentation suggests two other thinking mode controls: First: append `/think` or `/no_think` to user input for agent-level control. ```python messages = [ {"role": "user", "content": "Give me a short introduction to large language models./no_think"}, ] messages = generator(messages, max_new_tokens=32768)[0]["generated_text"] # print(messages[-1]["content"]) messages.append({"role": "user", "content": "In a single sentence./think"}) messages = generator(messages, max_new_tokens=32768)[0]["generated_text"] # print(messages[-1]["content"]) ``` Testing shows only `/no_think` works; `/think` doesn't. Second: append assistant-role message "`\n\n\n\n`" after each user input to temporarily disable thinking mode. ```python messages = [ {"role": "user", "content": "Give me a short introduction to large language models."}, {"role": "assistant", "content": "\n\n\n\n"}, ] messages = generator(messages, max_new_tokens=32768)[0]["generated_text"] # print(messages[-1]["content"]) messages.append({"role": "user", "content": "In a single sentence."}) messages = generator(messages, max_new_tokens=32768)[0]["generated_text"] # print(messages[-1]["content"]) ``` Testing shows this doesn't work. Thus, the most reliable method remains adding `extra_body` during `model_client` initialization. ### Other extra\_body options `extra_body` controls other Qwen3 features too: **top\_k:** Controls sampling candidate set size during generation. Configure via extra\_body={"top\_k":xxx}. **thinking\_budget:** Maximum thinking length, only effective when `enable_thinking=True`. Configure via `extra_body={"thinking_budget": xxx}` **translation\_options:** For translation models, configures source/target languages, e.g., `extra_body={"translation_options": { "source_lang": "auto", "target_lang": "English" }}`. Use "auto" for mixed languages. **enable\_search:** Whether to reference web searches before generation. Configure via `extra_body={"enable_search": True}` --- ## Practice Exercise: Article-Summarizing Agent Having learned to connect AutoGen with Qwen3, support structured\_output, and control thinking mode, let's test our knowledge with a small exercise. We'll use Qwen3 and mcp fetch server to create an agent that summarizes online articles with structured output. ![The business flow diagram of this agent practice.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/autogen_with_qwen3-practics-diagram.drawio.png) The business flow diagram of this agent practice. Image by Author First, initialize a `model_client` using `qwen3-30b-a3b` with thinking mode disabled: ```python model_client = OpenAILikeChatCompletionClient( model='qwen3-30b-a3b', temperature=0.01, extra_body={"enable_thinking": False} ) ``` For structured output, we'll define a data class extracting `title`, `url`, `author`, `keywords`, and `summary` from articles, adding descriptions for clarity: ```python class ArticleDetail(BaseModel): title: str url: str author: str = Field(..., description="The author of the article.") keywords: list[str] = Field(..., description="You need to provide me with no more than 5 keywords.") summary: str = Field(..., description=""" High level summary of the article with relevant facts and details. Include all relevant information to provide full picture. """) ``` Next, we'll run `mcp-server-fetch` locally, connecting via AutoGen's `StdioServerParams` and `StdioMcpToolAdapter` in main: ```python server_params = StdioServerParams( command="python", args=["-m", "mcp_server_fetch"], read_timeout_seconds=30 ) fetch = await StdioMcpToolAdapter.from_server_params(server_params, "fetch") ``` Now define the agent, passing `fetch` mcp tool and our data class. Note: if thinking mode is enabled, `model_client_stream` must be `True` for streaming output. Here we've disabled thinking mode but kept `model_client_stream=True`: ```python agent = AssistantAgent( name="web_browser", model_client=model_client, tools=[fetch], system_message="You are a helpful assistant.", output_content_type=ArticleDetail, model_client_stream=True ) ``` Finally, have the agent read my previous article and produce structured output: ```python result = await Console( agent.run_stream(task=""" Please visit https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/ and give me a quick summary. """) ) output = cast(ArticleDetail, result.messages[-1].content) console.print(Markdown(dedent(f""" \n \n **📋Title:** {output.title} **🔗Link:** {output.url} **🧑‍💻Author:** {output.author} **🏷️Keywords:** {output.keywords} **📃Summary:** {output.summary} """))) ``` Excellent! The agent successfully fetched the article and generated structured output via `ArticleDetail`: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/05/image-15.png) The agent successfully fetched the article and generated structured output. Image by Author Our `OpenAILikeChatCompletionClient` implementation hides all structured\_output and thinking mode details, making project code beautifully simple. --- ## Conclusion As Qwen3 deployment grows in enterprises, corresponding agent frameworks must keep pace. This article explored adapting AutoGen for Qwen3. Through various approaches, I've explained structured\_output principles that benefit you even beyond AutoGen. I've also analyzed AutoGen's parameter handling at code level, adding support for Qwen3's `extra_body` parameters. Having encapsulated these details in `OpenAILikeChatCompletionClient`, you'll find Qwen3 integration remarkably straightforward. In upcoming articles, I'll continue guiding you through modern multi-agent frameworks. Feel free to leave comments with questions. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- ## Recommended Enterprise AI Courses Getting the hang of some AI coding tricks is just the beginning. To learn the AI development tech stack more systematically and grow from an "AI enthusiast" to an "AI-driven software engineer," here are some courses I'd recommend to you: - Dive deeper into AI coding and agent building. Check out Vanderbilt University's [**Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) on Coursera. - To build real, production-ready enterprise AI agent apps to kickstart your AI career. The [**IBM RAG and Agentic AI Certificate**](https://imp.i384100.net/VOVRV3?ref=dataleadsfuture.com) is a great fit here. Both courses are included in [**Coursera Plus**](https://imp.i384100.net/3kOjjB?ref=dataleadsfuture.com). If you're planning to level up in both AI engineering and architecture this year, the annual subscription is definitely the better deal. --- ## Further Reading Give my programming workflow built with OpenCode, OMO-Slim, and OpenSpec a try? [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-37afe0dd-2c1f-48e6-8d66-1c39a54908fd.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-9ee163b8-2664-4664-8e83-6ac486661930.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) Can't your DeepSeek-V4 and GLM-5.2 agents read images yet? Give my method a try: [DeepSeek-V4 Can’t Read Images? I Made It ReadDon’t wait for a multimodal model, you can use it now![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-b60cae59-312e-4dc1-994b-4abee46cac8c.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-01b85923-e443-492f-84e3-c18ea2db8230.webp)](https://www.dataleadsfuture.com/deepseek-v4-cant-read-images-i-made-it-read/) The concept of Loop Engineering has been getting a lot of buzz lately, so I decided to give it a shot in OpenCode. The results were surprisingly good: [No Plugins Needed, I Built a Fully Automated Coding Loop in OpenCodeUsing DeepSeek-V4 for low-cost Loop Engineering![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-6f57293c-9130-4997-ae4b-ea4ec23b20cc.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-3-comprass-c968be2b-0904-497f-817b-9b6f740a5943.webp)](https://www.dataleadsfuture.com/no-plugins-needed-i-built-a-fully-automated-coding-loop-in-opencode/) --- ## Source Code Here's the source code for this article. Subscribe now to get it, 100% free to use however you like: [Subscribe Now ](#/portal/) ### Fixing the Agent Handoff Problem in LlamaIndex's AgentWorkflow System URL: https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/ Last updated: 2026-04-22T08:26:20.000Z [LlamaIndex AgentWorkflow](https://docs.llamaindex.ai/en/stable/understanding/agent/multi%5Fagent/?ref=dataleadsfuture.com), as a brand-new multi-agent orchestration framework, still has some shortcomings. The most significant issue is that after an agent hands off control, the receiving agent fails to continue responding to user requests, causing the workflow to halt. In today's article, I'll explore several experimental solutions to this problem with you and discuss the root cause behind it: the positional bias issue in LLMs. I've included all relevant source code at the end of this article. Feel free to read or modify it without needing my permission. --- ## Introduction My team and I have been experimenting with LlamaIndex AgentWorkflow recently. After some localization adaptations, we hope this framework can eventually run in our production system. During the adaptation, we encountered many obstacles. I've documented these problem-solving experiences in [my article series](https://www.dataleadsfuture.com/tag/agentic-ai/). You might want to read them first to understand the full context. Today, I'll address the issue where after the on-duty agent hands off control to the next agent, the receiving agent fails to continue responding to the user's most recent request. Here's what happens: ![The receiving agent doesn't immediately respond to the user's latest request - the user has to repeat their question. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image.png) The receiving agent doesn't immediately respond to the user's latest request - the user has to repeat their question. Image by Author After the handoff, the receiving agent doesn't immediately respond to the user's latest request - the user has to repeat their question. ### Why should I care? In this article, we'll examine this unique phenomenon and attempt to solve it from multiple perspectives, including developer recommendations and our own experience. During this process, we'll intentionally review AgentWorkflow's excellent source code, having a cross-temporal conversation with its authors through code to better understand Agentic AI design principles. We'll also touch upon LLM position bias for the first time, understanding how position bias in chat history affects LLM responses. These insights aren't limited to LlamaIndex - they'll help us handle similar situations when working with other multi-agent orchestration frameworks. Let's go. --- ## The Developer-Recommended Solution ### First, let's see what the developers say Before we begin, if you need background on LlamaIndex AgentWorkflow, feel free to read my previous article: [Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration SolutionAnd fix the issue where the agent can’t continue with past requests![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-25.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4-1-3.webp)](https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/) In short, LlamaIndex AgentWorkflow builds upon the excellent LlamaIndex Workflow framework, encapsulating agent function calling, handoff, and other cutting-edge Agentic AI developments. It lets you focus solely on your agent's business logic. In my previous article, I first mentioned the issue where agents fail to continue processing user requests after handoff. Others have noticed this too. In this thread, someone referenced my article's solution when asking the developers about it. I'm glad I could help: [I was building a multi-agent workflow, where each agent...@kapa.ai I was building a multi-agent workflow, where each agent has multiple tools. I started the workflow with user query, the query goes to the root agent an…![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/4dac85e2-45e7-4335-8cda-5c678dec0e57.png)LlamaIndexZenitsu![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/700176e5-9f26-4585-bdc5-78c3f47d3175)](https://community.llamaindex.ai/i-was-building-a-multiagent-workflow-where-each-agent-has-multiple-tools-i-started-the-FjYDXGDDCG2W?ref=dataleadsfuture.com) Developer Logan M proposed including the original user request in the `handoff` method's output to ensure the receiving agent continues processing. Unfortunately, as of this writing, LlamaIndex's release version hasn't incorporated this solution yet. So today's article starts with the developer's response - we'll try rewriting the `handoff` method implementation ourselves to include the original user request in the `handoff` output. ### First attempt Since this solution modifies the `handoff` method implementation, we don't need to rewrite `FunctionAgent` code. Instead, we'll modify `AgentWorkflow's` implementation. The `handoff` method is core to AgentWorkflow's handoff capability. It identifies which agent the LLM wants to hand off to and sets it in the context's `next_agent`. During workflow execution, this method merges with the agent's tools and gets called via function calling when the LLM needs to hand off. This is how AgentWorkflow implements multi-agent handoff. In the original code, after handoff sets the `next_agent`, it returns a prompt as the tool call result to the receiving agent. The prompt looks like this: ```python DEFAULT_HANDOFF_OUTPUT_PROMPT = """ Agent {to_agent} is now handling the request due to the following reason: {reason}. Please continue with the current request. """ ``` This prompt includes `{to_agent}` and `{reason}` fields. But since the prompt goes to the receiving agent, `{to_agent}` isn't very useful. Unless `{reason}` contains the original user request, the receiving agent can't get relevant information from this prompt. That's why the developer suggested including the user request in the prompt output. ![The original handoff implementation didn't include user request information.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/Fixing_AgentWorkflow_Problem-original_handoff.drawio.png) The original handoff implementation didn't include user request information. Image by Author Let's modify this method first. We'll create an `enhanced_agent_workflow.py` file and write the modified `HANDOFF_OUTPUT_PROMPT`: ```python ENHANCED_HANDOFF_OUTPUT_PROMPT = """ Agent {to_agent} is now handling the request. Check the previous chat history and continue responding to the user's request: {user_request}. """ ``` Compared to the original, I added a requirement for the LLM to review chat history and included the user's most recent request. ![The output of the updated handoff now includes both chat history review and user request information.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/Fixing_AgentWorkflow_Problem-modified_handoff.drawio.png) The output of the updated handoff now includes both chat history review and user request information. Image by Author Next, I rewrote the `handoff` method to return the new prompt: ```python async def handoff(ctx: Context, to_agent: str, user_request: str): """Handoff control of that chat to the given agent.""" agents: list[str] = await ctx.get('agents') current_agent_name: str = await ctx.get("current_agent_name") if to_agent not in agents: valid_agents = ", ".join([x for x in agents if x != current_agent_name]) return f"Agent {to_agent} not found. Please select a valid agent to hand off to. Valid agents: {valid_agents}" await ctx.set("next_agent", to_agent) handoff_output_prompt = PromptTemplate(ENHANCED_HANDOFF_OUTPUT_PROMPT) return handoff_output_prompt.format(to_agent=to_agent, user_request=user_request) ``` The rewrite is simple - I just changed the `reason` parameter to `user_request` and returned the new prompt. The LLM will handle everything else. Since we modified `handoff's` source code, we also need to modify `AgentWorkflow's` code that calls this method. The `_get_handoff_tool` method in `AgentWorkflow` calls `handoff`, so we'll implement an `EnhancedAgentWorkflow` subclass of `AgentWorkflow` and override `_get_handoff_tool`: ```python class EnhancedAgentWorkflow(AgentWorkflow): def _get_handoff_tool( self, current_agent: BaseWorkflowAgent ) -> Optional[AsyncBaseTool]: """Creates a handoff tool for the given agent.""" agent_info = {cfg.name: cfg.description for cfg in self.agents.values()} configs_to_remove = [] for name in agent_info: if name == current_agent.name: configs_to_remove.append(name) elif ( current_agent.can_handoff_to is not None and name not in current_agent.can_handoff_to ): configs_to_remove.append(name) for name in configs_to_remove: agent_info.pop(name) if not agent_info: return None handoff_prompt = PromptTemplate(ENHANCED_HANDOFF_PROMPT) fn_tool_prompt = handoff_prompt.format(agent_info=str(agent_info)) return FunctionTool.from_defaults( async_fn=handoff, description=fn_tool_prompt, return_direct=True ) ``` Our modifications are complete. Now let's write test code in `example_2.py` to verify our changes. (`example_1.py` contains the original AgentWorkflow test.) I'll base the code on this user's scenario to recreate the situation. We'll create two agents: `search_agent` and `research_agent`. `search_agent` searches the web and records notes, then hands off to `research_agent`, who writes a research report based on the notes. `search_agent`: ```python search_agent = FunctionAgent( name="SearchAgent", description="You are a helpful search assistant.", system_prompt=""" You're a helpful search assistant. First, you'll look up notes online related to the given topic and recorde these notes on the topic. Once the notes are recorded, you should hand over control to the ResearchAgent. """, tools=[search_web, record_notes], llm=llm, can_handoff_to=["ResearchAgent"] ) ``` `research_agent`: ```python research_agent = FunctionAgent( name="ResearchAgent", description="You are a helpful research assistant.", system_prompt=""" You're a helpful search assistant. First, you'll look up notes online related to the given topic and recorde these notes on the topic. Once the notes are recorded, you should hand over control to the ResearchAgent. """, llm=llm ) ``` `search_agent` is a multi-tool agent that uses `search_web` and `record_notes` methods: `search_web`: ```python async def search_web(ctx: Context, query: str) -> str: """ This tool searches the internet and returns the search results. :param query: user's original request :return: Then return the search results. """ tavily_client = AsyncTavilyClient() search_result = await tavily_client.search(str(query)) return str(search_result) ``` `record_notes`: ```python async def record_notes(ctx: Context, notes: str, notes_title: str) -> str: """ Useful for recording notes on a given topic. Your input should be notes with a title to save the notes under. """ return f"{notes_title} : {notes}" ``` Finally, we'll use `EnhancedAgentWorkflow` to create a workflow and test our modifications: ```python workflow = EnhancedAgentWorkflow( agents=[search_agent, research_agent], root_agent=search_agent.name ) async def main(): handler = workflow.run(user_msg="What is LLamaIndex AgentWorkflow, and what problems does it solve?") async for event in handler.stream_events(): if isinstance(event, AgentOutput): print("=" * 70) print(f"🤖 {event.current_agent_name}") if event.response.content: console.print(Markdown(event.response.content or "")) else: console.print(event.tool_calls) if __name__ == "__main__": asyncio.run(main()) ``` ![After research_agent takes over, it recognizes the user request but still doesn't respond. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-1.png) After research\_agent takes over, it recognizes the user request but still doesn't respond. Image by Author After `research_agent` takes over, it recognizes the user request but still doesn't respond. Our attempt failed. 😭 --- ## My Proposed Solution ### How I view this issue In my previous article, I speculated about the cause: ![FunctionAgent puts both tool_call and tool_call_result info into ChatMemory, which pushes user requests to the back of the queue.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-2.png) FunctionAgent puts both tool\_call and tool\_call\_result info into ChatMemory, which pushes user requests to the back of the queue. Image by Author As shown, `FunctionAgent` stores all chat messages in a `MemoryBuffer` \- essentially a FIFO queue where user requests enter first. After completing function calling based on user requests, `FunctionAgent` saves both `tool_call` and `tool_call_result` as chat messages in memory. Each function call generates two messages. Multiple tool calls create even more messages. This pushes the original user request deeper into the queue - either far from the latest message or, due to `MemoryBuffer's`After token limit, completely out of the queue. Consequently, the LLM struggles to perceive the original request from chat history. I'll explain the technical reasons in the position bias section. When the next agent takes over, it can't immediately respond to the user request. So I tried a simple fix: After each handoff, I copy the original user request to the queue's end, ensuring the LLM notices it. ![After each handoff, I copy the original user request to the queue's end.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/Fixing_AgentWorkflow_Problem.second_attemp.drawio.png) After each handoff, I copy the original user request to the queue's end. Image by Author ### Second attempt This attempt's code is in `reordered_function_agent.py`. The implementation is simple: I subclass `FunctionAgent` as `ReorderedFunctionAgent` and override `take_step`. ```python class ReorderedFunctionAgent(FunctionAgent): @override async def take_step( self, ctx: Context, llm_input: List[ChatMessage], tools: Sequence[AsyncBaseTool], memory: BaseMemory, ) -> AgentOutput: last_msg = llm_input[-1] and llm_input[-1].content state = await ctx.get("state", None) if "handoff_result" in last_msg: for message in llm_input[::-1]: if message.role == MessageRole.USER: last_user_msg = message llm_input.append(last_user_msg) break return await super().take_step(ctx, llm_input, tools, memory) ``` When I detect the last message in `llm_input` is a handoff `tool_call_result`, I traverse backward to find the user's last request and append it to the queue's end. To identify handoff `tool_call_result` messages, I manually pass a `handoff_output_prompt` during `AgentWorkflow` initialization, adding a "handoff\_result:" string as a marker. The test code is in `example_3.py`: ```python workflow = AgentWorkflow( agents=[search_agent, research_agent], root_agent=search_agent.name, handoff_output_prompt=( "handoff_result: Due to {reason}, the user's request has been passed to {to_agent}." "Please review the conversation history immediately and continue responding to the user's request." ), ) ``` Let's run the test: ![This time, research_agent's output is just a summary of partial notes rather than a final research report.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-3.png) This time, research\_agent's output is just a summary of partial notes rather than a final research report. Image by Author This time, `research_agent `successfully detects and responds to the user request after taking over. But the result isn't perfect - it doesn't realize the web search and note-taking already happened. It thinks research isn't complete, so the output is just a summary of partial notes rather than a final research report. I believe this happens because after appending the user request to `ChatMemory's` end, previous `tool_call` information gets pushed to the front, causing the LLM to lose critical information. Next, we'll examine the theoretical basis of this problem and propose an ultimate solution. --- ## Theoretical Cause: Position Bias of LLMs This issue of messages at the queue's front being ignored relates to a rarely discussed topic: position bias. Since this isn't an academic discussion, I won't cite many research papers or delve deep into theory. If interested, search for "position bias of large language model." I'll explain this phenomenon in simple terms: ![Different positions in the chat context have different attention weights.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-4.png) Different positions in the chat context have different attention weights. [Arxiv 2407.01100](https://arxiv.org/pdf/2407.01100?ref=dataleadsfuture.com) Our text instructions to LLMs typically include two segments: `system_prompt`an average and chat history - collectively called the LLM's context. LLMs have an attention weight decay mechanism. As context expands, attention weights for earlier information naturally decay. When knowledge sits at chat history's front, its influence diminishes rapidly with new dialogue turns. Experiments show that in an 8k token context window, tokens in the first 10% positions see an average over 60% influence weight drop. [(Large Language Model Agent: A Survey on Methodology, Applications and Challenges - Junyu Luo et al., 2025)](https://arxiv.org/pdf/2503.21460?ref=dataleadsfuture.com) System prompts are designed as global control signals, with information there having higher confidence (about 3- 5x weight difference). ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-5.png) In different LLMs, the positions where the model focuses on important info don't always match the actual important info spots. [Arxiv 2404.01430](https://arxiv.org/pdf/2404.01430?ref=dataleadsfuture.com) Imagine entering a restaurant. You first notice the menu cover (system prompt) featuring special dishes and chef introductions, then seasonal items (latest chat history), and finally regular dishes. Delicacies hidden in regular dishes often get overlooked. ![Just like how people always notice the menu cover and seasonal specials.](https://images.unsplash.com/photo-1627907228175-2bf846a303b4?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=M3wxMTc3M3wwfDF8c2VhcmNofDEyMnx8bWVudXxlbnwwfHx8fDE3NDQwODEzODZ8MA&ixlib=rb-4.0.3&q=80&w=2000) Just like how people always notice the menu cover and seasonal specials. Photo by [Frank Holleman](https://unsplash.com/@fraenkly?ref=dataleadsfuture.com) / [Unsplash](https://unsplash.com/?utm%5Fsource=ghost&utm%5Fmedium=referral&utm%5Fcampaign=api-credit) Understanding the cause leads us to the ultimate solution. --- ## My Final Attempt ### What I plan to do Next, I'll walk you through my final attempt. First, here's what our project output looks like after implementing it: ![ResearchAgent not only continues processing the user request but fully perceives the search notes, ultimately producing a perfect research report.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-6.png) ResearchAgent not only continues processing the user request but fully perceives the search notes, ultimately producing a perfect research report. Image by Author After taking over, `ResearchAgent` not only continues processing the user request but fully perceives the search notes, ultimately producing a perfect research report. My solution approach: I'll discard `tool_call` and `tool_call_result` information, no longer appending them to chat history. Chat history will only keep user requests (role: user) and LLM outputs (role: assistant). Where should we put external information from `tool_call`? Here we'll use another AgentWorkflow feature: the original framework lets you initialize a state in `Context` for storing predefined system states with persistence support. But in the original framework, state information accompanies each user request. This limits state usage scenarios - it mostly stores static content like login information. So I'll modify this. I'll place useful information from `tool_call` in state information, no longer including it in user requests. Instead, I'll put state information in `system_prompt`. Remember I mentioned `system_prompt` information has about 3-5x higher confidence weight? If we want the LLM to notice important information, `system_prompt `is ideal. This aligns with my project experience. In a previous project using [vanna.ai](https://vanna.ai/?ref=dataleadsfuture.com), we initially placed few-shot prompting examples in chat history, resulting in low accuracy. After moving few-shot examples to `system_prompt`, LLM generation accuracy improved dramatically. Try it yourself. ![How the placement of information affects the accuracy of LLMs.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/Fixing_AgentWorkflow_Problem-accuracy.drawio.png) How the placement of information affects the accuracy of LLMs. Image by Author Back to today's article - I'll make similar adjustments to AgentWorkflow: keeping only user-LLM chat messages in chat history, while placing all external information from `tool_call` into agent `system_prompt` via state. How? Let's proceed to implementation. ### Code implementation To implement this solution, I need to modify both `AgentWorkflow` and `FunctionAgent` classes. In AgentWorkflow, I'll remove the inclusion of state information in user requests. We'll create `contextual_agent_workflow.py` with a `ContextualAgentWorkflow` subclass of `AgentWorkflow`. In `ContextualAgentWorkflow`, we override `init_run` to simply place user requests in `ChatMemory` without state information: ```python class ContextualAgentWorkflow(AgentWorkflow): @step async def init_run(self, ctx: Context, ev: StartEvent) -> AgentInput: """Sets up the workflow and validates inputs""" await self._init_context(ctx, ev) user_msg: Optional[Union[str, ChatMessage]] = ev.get("user_msg") chat_history: Optional[List[ChatMessage]] = ev.get("chat_history", []) memory: BaseMemory = await ctx.get("memory") current_agent_name: str = await ctx.get("current_agent_name") if isinstance(user_msg, str): user_msg = ChatMessage(role="user", content=user_msg) if user_msg: await memory.aput(user_msg) await ctx.set("user_msg_str", user_msg.content) elif chat_history: last_msg = chat_history[-1].content or "" memory.set(chat_history) await ctx.set("user_msg_str", last_msg) else: raise ValueError("Must provide either user_msg or chat_history") input_messages = memory.get() return AgentInput(input=input_messages, current_agent_name=current_agent_name) ``` After modifying `AgentWorkflow`, we'll adjust `FunctionAgent`. Create `contextual_function_agent.py` with a `ContextualFunctionAgent` subclass of `FunctionAgent`. In `contextual_function_agent.py`, we'll add a new `STATE_STR_PROMPT` string containing state information that will ultimately append to `system_prompt`: ```python STATE_STR_PROMPT = """ Current state: {state_str} """ ``` We'll also keep an option to modify the default `STATE_STR_PROMPT` in `ContextualFunctionAgent`: ```python class ContextualFunctionAgent(FunctionAgent): """The Function Agent contains a system_prompt with state strings.""" state_str_prompt: Optional[str] = Field( default=STATE_STR_PROMPT, description="Adding state information to the system_prompt." ) ``` Next, we'll override `FunctionAgent` methods. `FunctionAgent` implements `take_step`, `handle_tool_call_results`, and `finalize`. I'll override `take_step` and `handle_tool_call_results`. ![Attach the tool call result as state info in the system_prompt.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/Fixing_AgentWorkflow_Problem-contextual.drawio.png) Attach the tool call result as state info in the system\_prompt. Image by Author In `take_step`, I removed statements saving `tool_call` information to `Context`, built `STATE_STR_PROMPT`, appended it to `system_prompt`, then normally called the parent class's `take_step` to get LLM results. ```python class ContextualFunctionAgent(FunctionAgent): """The Function Agent contains a system_prompt with state strings.""" ... @override async def take_step( self, ctx: Context, llm_input: List[ChatMessage], tools: Sequence[AsyncBaseTool], memory: BaseMemory ) -> AgentOutput: if '{state_str}' not in self.state_str_prompt: raise ValueError("{state_str} not found in provided state_str_prompt") current_state = await ctx.get("state") state_str_template = PromptTemplate(self.state_str_prompt) state_prompt = state_str_template.format( state_str=current_state ) if llm_input[0].role == "system": llm_input[0].content = llm_input[0].content + state_prompt else: llm_input = [ChatMessage(role="system", content=state_prompt)] + llm_input output = await super().take_step( ctx, llm_input, tools, memory ) await ctx.set(self.scratchpad_key, []) return output ``` `handle_tool_call_results` doesn't call parent methods. Its key difference is not writing `tool_call_result` to `ChatMemory`, but recording it in Context's state. During the next `take_step` call, the latest state information appends to `system_prompt`. ```python class ContextualFunctionAgent(FunctionAgent): """The Function Agent contains a system_prompt with state strings.""" ... @override async def handle_tool_call_results( self, ctx: Context, results: List[ToolCallResult], memory: BaseMemory ) -> None: current_state = await ctx.get("state", {}) for tool_call_result in results: if ( tool_call_result.return_direct and tool_call_result.tool_name != "handoff" ): await memory.aput( ChatMessage( role="assistant", content=str(tool_call_result.tool_output.content), additional_kwargs={"tool_call_id": tool_call_result.tool_call_id} ) ) break current_state[tool_call_result.tool_name] = str(tool_call_result.tool_output.content) await ctx.set("state", current_state) ``` Our modifications to `AgentWorkflow` and `FunctionAgent` are complete. Now let's modify the test code. The modification is simple - just replace original `FunctionAgent` and `AgentWorkflow` with `ContextualFunctionAgent` and `ContextualAgentWorkflow`. ```python search_agent = ContextualFunctionAgent( name="SearchAgent", description="You are a helpful search assistant.", ... ) research_agent = ContextualFunctionAgent( name="ResearchAgent", description="You are a helpful research assistant.", ... ) workflow = ContextualAgentWorkflow( agents=[search_agent, research_agent], root_agent=search_agent.name ) ``` Running this code gives excellent results: ![We've indeed identified and effectively solved the root problem.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/04/image-7.png) We've indeed identified and effectively solved the root problem. Image by Author We've indeed identified and effectively solved the root problem. --- ## Conclusion LlamaIndex AgentWorkflow is a great multi-agent programming framework, but it still has some flaws: when the on-duty agent hands over control to the next agent, the receiving agent sometimes fails to continue responding to user requests. In today's article, we first tried the official recommended method and the approach I proposed in my previous article to solve this issue, but the problem couldn't be effectively resolved. So, I explored a relatively niche topic with you: LLM's position bias. We learned that LLMs assign different parameter weights to information in different positions of the context. Based on this technical theory, we attempted a temporary modification method and succeeded. However, the real world is more complex than experiments. Besides position bias, factors like the user's system\_prompt and different LLMs' preferences for information positions also contribute to this issue. This requires LlamaIndex to address it at the framework level. I'm still waiting for LlamaIndex AgentWorkflow to become a truly enterprise-ready, production-grade multi-agent orchestration framework. Until then, I'll pause this magical journey through the Workflow series. But I look forward to meeting LlamaIndex AgentWorkflow again in the near future. Go LlamaIndex! --- The source code mentioned in this article is available here. Feel free to read and modify it without needing my permission: [agentic-ai-playground/02\_Fix\_LlamaIndex\_AgentWorkflow\_Couldnot\_Continue at main · qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-3.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground-1)](https://github.com/qtalen/agentic-ai-playground/tree/main/02%5FFix%5FLlamaIndex%5FAgentWorkflow%5FCouldnot%5FContinue?ref=dataleadsfuture.com) ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Integrating LlamaIndex and DeepSeek-R1 for reasoning_content and Function Calling Features URL: https://www.dataleadsfuture.com/integrating-llamaindex-and-deepseek-r1-for-reasoning_content-and-function-call-features-2/ Last updated: 2026-04-22T08:30:22.000Z In today's article, I'll show you how to use LlamaIndex's AgentWorkflow to read the reasoning process from DeepSeek-R1's output and how to enable function calling features for DeepSeek-R1 within AgentWorkflow. All the source code discussed is available at the end of this article for you to read and modify freely. --- ## Introduction DeepSeek-R1 is incredibly useful. Most of my work is now done with its help, and major cloud providers support its deployment, making API access even easier. Most open-source models support OpenAI-compatible API interfaces, and DeepSeek-R1 is no exception. However, unlike other generative models, DeepSeek-R1 includes a `reasoning_content` key in its output, which stores the Chain-of-Thought (CoT) reasoning process. Additionally, DeepSeek-R1 doesn't support function calling, which makes developing agents with it quite challenging. (The official documentation claims that DeepSeek-R1 doesn't support structured output, but in my tests, it does. So, I won't cover that here.) ### Why should I care? While you can find a DeepSeek client on llamahub, it only supports DeepSeek-V3 and doesn't read the `reasoning_content` from DeepSeek-R1. Moreover, LlamaIndex is now fully focused on building AgentWorkflow. This means that even if you only use the LLM as the final step in your workflow for content generation, it still needs to support handoff tool calls. ![Every agent should have the ability to hand off control to another agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-11.png) Every agent should have the ability to hand off control to another agent. Image by Author In this article, I'll show you two different ways to extend LlamaIndex to read `reasoning_content`. We'll also figure out how to make DeepSeek-R1 support function calling so it can be used in AgentWorkflow. Let's get started. --- ## Modifying OpenAILike Module for reasoning\_content Support ### Building a DeepSeek client To allow AgentWorkflow to read `reasoning_content`, we first need to make LlamaIndex's model client capable of reading this key. Let's start with the client module and see how to read `reasoning_content`. Previously, I wrote about connecting LlamaIndex to privately deployed models: [How to Connect LlamaIndex with Private LLM API DeploymentsWhen your enterprise doesn’t use public models like OpenAI![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-22.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LlamaIndex_LLM_API.drawio-1.png)](https://www.dataleadsfuture.com/how-to-connect-llamaindex-with-private-llm-api-deployments/) In short, you can't directly use LlamaIndex's OpenAI module to connect to privately or cloud-deployed large language model APIs, as it restricts you to GPT models. OpenAILike extends the OpenAI module by removing model type restrictions, allowing you to connect to any open-source model. We'll use this module to create our DeepSeek-R1 client. OpenAILike supports several methods, like `complete`, `predict`, and `chat`. Considering DeepSeek-R1's CoT reasoning, we often use the `chat` method with stream output. We'll focus on extending the chat scenario today. Since `reasoning_content` appears in the model's output, we need to find methods handling model output. In OpenAILike, `_stream_chat` and `_astream_chat` handle sync and async output, respectively. We can create a DeepSeek client module, inheriting from OpenAILike, and override the `_stream_chat` and `_astream_chat` methods. Here's the code skeleton: ```python class DeepSeek(OpenAILike): @llm_retry_decorator def _stream_chat( self, messages: Sequence[ChatMessage], **kwargs: Any ) -> ChatResponseGen: ... @llm_retry_decorator async def _astream_chat( self, messages: Sequence[ChatMessage], **kwargs: Any ) -> ChatResponseAsyncGen: ... ``` We'll also need to inherit `ChatResponse` and create a `ReasoningChatResponse` model to store the `reasoning_content` we read: ```python class ReasoningChatResponse(ChatResponse): reasoning_content: Optional[str] = None ``` Next, we'll implement the `_stream_chat` and `_astream_chat` methods. In the parent class, these methods are similar, using closures to build a Generator for structuring streamed output into a `ChatResponse` model. The `ChatResponse` model contains parsed `message`, `delta`, and `raw` information. Since OpenAILike doesn't parse `reasoning_content`, we'll extract this key from raw later. We can first call the parent method to get the Generator, then build our own closure method: ```python class DeepSeek(OpenAILike): @llm_retry_decorator def _stream_chat( self, messages: Sequence[ChatMessage], **kwargs: Any ) -> ChatResponseGen: responses = super()._stream_chat(messages, **kwargs) def gen() -> ChatResponseGen: for response in responses: if processed := self._build_reasoning_response(response): yield processed return gen() ``` Since both methods handle processing similarly, we'll create a `_build_reasoning_response` method to handle `response.raw`. In streamed output, `reasoning_content` is found in `delta`, at the same level as content, as incremental output: ```python class DeepSeek(OpenAILike): ... @staticmethod def _build_reasoning_response(response: ChatResponse) \ -> ReasoningChatResponse | None: if not (raw := response.raw).choices: return None if (delta := raw.choices[0].delta) is None: return None return ReasoningChatResponse( message=response.message, delta=response.delta, raw=response.raw, reasoning_content=getattr(delta, "reasoning_content", None), additional_kwargs=response.additional_kwargs ) ``` With this, our DeepSeek client supporting `reasoning_content` streamed output is complete. The full code is available at the end of this article. Next, let's use Chainlit to create a simple chatbot to verify our modifications. ### Verifying reasoning\_content output with Chainlit Using Chainlit to display DeepSeek-R1's reasoning process is simple. Chainlit supports `step` objects, collapsible page blocks perfect for displaying reasoning processes. Here's the final result: ![The reasoning process of DeepSeek-R1.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-10.png) The reasoning process of DeepSeek-R1\. Image by Author First, we use `dotenv` to read environment variables, then use our `DeepSeek` class to build a client, ensuring `is_chat_model` is set to True: ```python load_dotenv("../.env") client = DeepSeek( model="deepseek-reasoner", is_chat_model=True, temperature=0.6 ) ``` Now, let's call the model and use `step` to continuously output the reasoning process. Since we're using a Generator to get `reasoning_content`, we can't use the `@step` decorator to call the `step` object. Instead, we use `async with` to create and use a step object: ```python @cl.on_message async def main(message: cl.Message): responses = await client.astream_chat( messages=[ ChatMessage(role="user", content=message.content) ] ) output = cl.Message(content="") async with cl.Step(name="Thinking", show_input=False) as current_step: async for response in responses: if (reasoning_content := response.reasoning_content) is not None: await current_step.stream_token(reasoning_content) else: await output.stream_token(response.delta) await output.send() ``` As shown in the previous result, our `DeepSeek` client successfully outputs `reasoning_content`, allowing us to integrate AgentWorkflow with DeepSeek-R1. --- ## Integrating AgentWorkflow and DeepSeek-R1 In my previous article, I detailed LlamaIndex's AgentWorkflow, an excellent multi-agent orchestration framework: [Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration SolutionAnd fix the issue where the agent can’t continue with past requests![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-23.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4-1-2.webp)](https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/) In that article, I mentioned that if the model supports function calling, we can use `FunctionAgent`; if not, we need to use `ReActAgent`. Today, I'll show you how to use LlamaIndex's `ReActAgent` to enable function calling support for the DeepSeek-R1 model. ### Enabling function calling with ReActAgent Why insist on function calling support? Can't DeepSeek just be a reasoning model for generating final answers? Unfortunately, no, because of the `handoff` tool. AgentWorkflow's standout feature is its agent handoff capability. When an agent decides it can't handle a user's request and needs to hand over control to another agent, it uses the handoff tool. This means an agent must support function calling to enable handoff, even if it doesn't need to call any tools. You might point out that the AgentWorkflow source code explains if an agent is configured with `can_handoff_to` but the list is empty, it won't concatenate the `handoff` tool, so no function calling is needed. ![When can_handoff_to is an empty list, the handoff_tool won't be concatenated.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-5.png) When `can_handoff_to` is an empty list, the `handoff_tool` won't be concatenated. Image by Author That's not the case. Under normal circumstances, we create agents based on FunctionAgent. FunctionAgent's `take_step` method calls the model client's `astream_chat_with_tools` and `get_tool_calls_from_response` methods, both related to function calling. If our model doesn't support function calling, it'll throw an error: ![When you use DeepSeek-R1 in FunctionAgent, it will throw an error.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-7.png) When you use DeepSeek-R1 in FunctionAgent, it will throw an error. Image by Author So what should we do? Luckily, besides `FunctionAgent`, AgentWorkflow offers `ReActAgent`, designed for non-function calling models. `ReActAgent` implements a "think-act-observe" loop, suitable for complex tasks. But today, our focus is on using `ReActAgent` to enable DeepSeek-R1's function calling feature. Let's demonstrate with a simple search agent script: First, we define a method for web search: ```python async def search_web(query: str) -> str: """ A tool for searching information on the internet. :param query: keywords to search :return: the information """ client = AsyncTavilyClient() return str(await client.search(query)) ``` Next, we configure an agent using `ReActAgent`, adding the `search_web` method to tools. Since we only have one agent in our workflow, we won't configure `can_handoff_to`: ```python search_agent = ReActAgent( name="SearchAgent", description="A helpful agent", system_prompt="You are a helpful assistant that can answer any questions", tools=[search_web], llm=llm ) ``` We'll set up AgentWorkflow with the agent and write a main method to test it: ```python workflow = AgentWorkflow( agents=[search_agent], root_agent=search_agent.name ) async def main(): handler = workflow.run(user_msg="If I use a $1000 budget to buy red roses in bulk for Valentine's Day in 2025, how much money can I expect to make?") async for event in handler.stream_events(): if isinstance(event, AgentStream): print(event.delta, end="", flush=True) ``` ![ReActAgent uses the search_web tool to fetch information.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-8.png) ReActAgent uses the search\_web tool to fetch information. Image by Author As seen, `ReActAgent` provides the method name `search_web` and parameters in Action and Action Input. The method executes successfully, indicating DeepSeek-R1 now supports function calling through `ReActAgent`. Hooray. ### Enabling reasoning\_content output with ReActAgent If you run the previous code, you'll notice that it takes a long time from execution to result return. This is because, as a reasoning model, DeepSeek-R1 undergoes a lengthy CoT reasoning process before generating the final result. But we extended our DeepSeek client for `reasoning_content`, and configured it in AgentWorkflow. So why no reasoning output? `ReActAgent` uses `AgentStream` for streaming model results. However, `AgentStream` lacks a `reasoning_content` attribute, preventing client-returned reasoning from being packaged in `AgentStream`. Should we modify `ReActAgent`'s source code, like we did with `OpenAILike`? Actually, no. Earlier, we modified `OpenAILike` because the model client is a foundational module, and we needed it to parse `reasoning_content` from raw text. But `AgentStream` has a `raw` property that includes the original message text, including `reasoning_content`. So, we just need to parse `reasoning_content` from `raw` after receiving the `AgentStream` message. ![We just need to parse the reasoning_content from the AgentStream.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/Integrating_LlamaIndex_and_DeepSeek-R1.ReasoningStreamAdapter.drawio.png) We just need to parse the reasoning\_content from the AgentStream. Image by Author We'll create a `ReasoningStreamAdapter` class, accepting an `AgentStream` instance in its constructor and exposing `AgentStream` attributes with `@property`, adding a `reasoning_content` attribute: ```python class ReasoningStreamAdapter: def __init__(self, event: AgentStream): self.event = event @property def delta(self) -> str: return self.event.delta @property def response(self) -> str: return self.event.response @property def tool_calls(self) -> list: return self.event.tool_calls @property def raw(self) -> dict: return self.event.raw @property def current_agent_name(self) -> str: return self.event.current_agent_name @property def reasoning_content(self) -> str | None: raw = self.event.raw if raw is None or (not raw['choices']): return None if (delta := raw['choices'][0]['delta']) is None: return None return delta.get("reasoning_content") ``` Usage is similar to this: ```python async for event in handler.stream_events(): if isinstance(event, AgentStream): adapter = ReasoningStreamAdapter(event) if adapter.reasoning_content: print(adapter.reasoning_content, end="", flush=True) if isinstance(event, AgentOutput): print(event) ``` ### Testing the final result with chat UI After modifying DeepSeek-R1 for function calling and `reasoning_content`, it's time to test our work. We'll use the Chainlit UI to observe our modifications: ![Now, AgentWorkflow can smoothly output the reasoning_content.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-9.png) Now, AgentWorkflow can smoothly output the reasoning\_content. Image by Author As shown, we ask the agent a complex task, and DeepSeek performs chain-of-thought reasoning while using the `search_web` tool to gather external information and return a solution. Due to space constraints, the complete project code isn't included here, but you'll find it at the end of this article. --- In the following story, this journey about LlamaIndex AgentWorkflow is coming to its final stop. We'll connect theory with practice and revisit the issue of AgentWorkflow not continuing after handoff: [Fixing the Agent Handoff Problem in LlamaIndex’s AgentWorkflow SystemThe position bias in LLMs is the root cause of the problem![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-27.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/p60Imf7mlg8tNL8kC9VdJSw8AsixrM-1.webp)](https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/) --- ## Further Reading: ReActAgent Implementation and Simplification If we review the earlier implementation, we'll notice that agents based on `ReActAgent` generate a lot of output. This is normal because `ReActAgent`'s reasoning-reflection mechanism requires multiple function calls to gather enough information to generate an answer. But our goal was simply to find a way to make DeepSeek-R1 support function calling. So, is there a way to simplify the agent's output? To achieve this, we should first understand how `ReActAgent` works. ![The parsing process of ReActAgent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/Integrating_LlamaIndex_and_DeepSeek-R1---3--.drawio.png) The parsing process of ReActAgent. Image by Author Reading `ReActAgent`'s source code, you'll find that implementing function calling involves two modules: `ReActChatFormatter` and `ReActOutputParser`. ### ReActChatFormatter ReActChatFormatter centers on a `system_prompt` satisfying the react architecture. The `system_prompt` has a `{tool_desc}` placeholder, replaced with current tools and descriptions at runtime. There's also a `{context_prompt}` placeholder, replaced with the ReActAgent's original `system_prompt`. When tool calls are needed, the `system_prompt` requires the model to return the following format: ```python ``` Thought: The current language of the user is: (user's language). I need to use a tool to help me answer the question. Action: tool name (one of {tool_names}) if using a tool. Action Input: the input to the tool, in a JSON format representing the kwargs (e.g. {{"input": "hello world", "num_beams": 5}}) ``` ``` Thought is reasoning, Action is tool name, Action Input is tool input, like `{"input": "hello world", "num_beams": 5}`. The reasoning process may iterate several times until the model gathers enough information to generate an answer, returning this: ```python ``` Thought: I can answer without using any more tools. I'll use the user's language to answer Answer: [your answer here (In the same language as the user's question)] ``` ``` If the model lacks information to answer, it returns: ```python ``` Thought: I cannot answer the question with the provided tools. Answer: [your answer here (In the same language as the user's question)] ``` ``` ### ReActOutputParser `ReActOutput` is simpler, parsing ReAct's step-by-step results. If it finds an `Action:` marker, it uses regex to extract and return the tool name and input. ```python pattern = ( r"\s*Thought: (.*?)\n+Action: ([^\n\(\) ]+).*?\n+Action Input: .*?(\{.*\})" ) match = re.search(pattern, input_text, re.DOTALL) ``` If there's no `Action:` marker but an `Answer:` marker, it extracts and returns the content after Answer: as the final result. ```python pattern = r"\s*Thought:(.*?)Answer:(.*?)(?:$)" match = re.search(pattern, input_text, re.DOTALL) ``` ![The way the parser works depends on what's in the output.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/Integrating_LlamaIndex_and_DeepSeek-R1-Page-2.drawio.png) The way the parser works depends on what's in the output. Image by Author Now we understand: - To make the model call a tool, it must generate content with Action and Action Input markers for tool name and input. - To avoid function calls, have the model generate content with an Answer: marker instead of `Action:`. Thus, we don't need to change `ReActOutputParser`, only pass a custom Formatter with a simplified `system_prompt`. Due to space constraints, I won't implement a simplified `ReActAgent`. Why not challenge yourself to try it? --- Here is the source code of this article. Feel free to read and modify it without needing my permission: [GitHub - qtalen/agentic-ai-playgroundContribute to qtalen/agentic-ai-playground development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-2.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/agentic-ai-playground)](https://github.com/qtalen/agentic-ai-playground?ref=dataleadsfuture.com) --- ## Conclusion DeepSeek-R1, a powerful reasoning model, adds a `reasoning_content` key to OpenAI API-compatible output for reasoning processes. However, its lack of function calling support limits its AgentWorkflow use. In this article, I made small LlamaIndex modifications to let users access `reasoning_content` via `AgentStream`, and combined `ReActAgent` with DeepSeek-R1 to enable function calling support. I also explored ways to simplify `ReActAgent`'s output by explaining its implementation. What are your thoughts on integrating LlamaIndex and DeepSeek? Feel free to comment, and I'll respond promptly. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ### Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration Solution URL: https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/ Last updated: 2026-07-06T10:04:24.000Z This article introduces you to the latest [AgentWorkflow](https://docs.llamaindex.ai/en/stable/understanding/agent/?ref=dataleadsfuture.com) multi-agent orchestration framework by LlamaIndex, demonstrating its application through a project, highlighting its drawbacks, and explaining how I solved them. By reading this, you'll learn how to simplify multi-agent orchestration and boost development efficiency using LlamaIndex AgentWorkflow. The project source code discussed here is available at the end of the article for your review and modification without my permission. --- ## Introduction Recently, I had to review LlamaIndex's official documentation for work and was surprised by the drastic changes: LlamaIndex has rebranded itself from a RAG framework to a multi-agent framework integrating data and workflow. [The entire documentation](https://docs.llamaindex.ai/en/stable/?ref=dataleadsfuture.com) is now built around AgentWorkflow. Multi-agent orchestration is not new. For enterprise-level applications, we don’t use a standalone agent to perform a series of tasks. Instead, we prefer a framework that can orchestrate multiple agents to collaborate on completing complex business scenarios. When it comes to multi-agent orchestration frameworks, you've probably heard of LangGraph, CrewAI, and AutoGen. However, LlamaIndex, once a framework as popular as LangChain, seemed silent in the multi-agent space in the past six months. Considering LlamaIndex’s high maturity and community involvement, the release of LlamaIndex AgentWorkflow caught our attention. So, my team and I studied it for a month and found that for practical applications, AgentWorkflow is a nearly perfect multi-agent orchestration solution. Smart as you might be, you might ask, since LlamaIndex Workflow has been out for half a year, what's the difference between Workflow and AgentWorkflow? To answer this, we must first look at how to use LlamaIndex Workflow for multi-agent setups. --- ## What Is Workflow? I previously wrote an article detailing what LlamaIndex Workflow is and how to use it: [Deep Dive into LlamaIndex Workflow: Event-driven LLM architectureWhat I think about the progress and shortcomings after practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LLamaIndex_Workflow-2.webp)](https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/) In simple terms, Workflow is an event-driven framework using Python asyncio for concurrent API calls to large language models and various tools. I also wrote about implementing multi-agent orchestration similar to OpenAI Swarm's agent handoff using Workflow: [Using LLamaIndex Workflow to Implement an Agent Handoff Feature Like OpenAI SwarmExample: a customer service chatbot project![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-1.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-1-2.webp)](https://www.dataleadsfuture.com/using-llamaindex-workflow-to-implement-an-agent-handoff-feature-like-openai-swarm/) However, Workflow is a relatively low-level framework and quite disconnected from other LlamaIndex modules, necessitating frequent learning and calls to LlamaIndex's underlying API when implementing complex multi-agent logic. If you’ve read my article, you'll notice I heavily rely on LlamaIndex’s low-level API across Workflow’s `step` methods for function calls and process control, leading to tight coupling between the workflow and agent-specific code. This isn’t ideal for those of us who want to finish work early and enjoy dinner at home. Perhaps LlamaIndex heard developers’ appeals, leading to the birth of AgentWorkflow. --- ## How Does AgentWorkflow Work? AgentWorkflow consists of an AgentWorkflow module and an Agent module. Unlike existing LlamaIndex modules, both are specially tailored for recent multi-agent objectives. Here, let’s first discuss the Agent module: ### Agent module The Agent module primarily consists of two classes: `FunctionAgent` and `ReActAgent`, both inheriting from `BaseWorkflowAgent`, hence incompatible with previous Agent classes. Use `FunctionAgent` if your language model supports function calls; if not, use `ReActAgent`. In this article, we use function calls to complete specific tasks, so we’ll focus on `FunctionAgent`: `FunctionAgent` mainly has three methods: `take_step`, `handle_tool_call_results`, and `finalize`. ![Illustrations of various methods in FunctionAgent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.functionagent.drawio.png) Illustrations of various methods in FunctionAgent. Image by Author The `take_step` method receives the current chat history `llm_input`, and available tools for the agent. It uses `astream_chat_with_tools` and `get_tool_calls_from_response` to get the next tools to execute, storing tool call parameters in the Context. Besides, `take_step` outputs the current round’s agent parameters and results in a stream, facilitating debugging and step-by-step viewing of intermediate agent execution results. The `handle_tool_call_results` method doesn’t directly execute tools – tools are invoked concurrently in AgentWorkflow. It merely saves tool execution results in the Context. The `finalize` method accepts an `AgentOutput` parameter but doesn’t alter it. Instead, it extracts tool call stacks from the Context, saving them as chat history in ChatMemory. You can inherit and override `FunctionAgent` methods to implement your business logic, which I’ll demonstrate in the upcoming project practice. ### Agentworkflow module Having covered the Agent module, let’s delve into the AgentWorkflow module. [In previous projects](https://www.dataleadsfuture.com/using-llamaindex-workflow-to-implement-an-agent-handoff-feature-like-openai-swarm/), I implemented an orchestration process based on Workflow. This was the flowchart at that time: ![The flowchart of the workflow implemented in the previous article.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image.png) The flowchart of the workflow implemented in the previous article. Image by Author Since my code referenced LlamaIndex's official examples, AgentWorkflow closely resembles my implementation but is simplified as it extracts the handoff and function call logic. Here’s AgentWorkflow’s architecture: ![The architecture diagram of AgentWorkflow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.AgentFlow.FlowChart.drawio.png) The architecture diagram of AgentWorkflow. Image by Author The entry point is the `init_run` method, which initializes Context and ChatMemory. Next, `setup_agent` identifies the duty agent, extracting its `system_prompt` and merging it with the current ChatHistory. Then, `run_agent_step` calls the agent’s `take_step` to obtain the required tools for invocation while writing large language model call results to the output stream. In the upcoming project practice, I’ll rewrite `take_step` for project-specific execution. Notably, `handoff`, incorporated as a tool, integrates into agent-executable tools within `run_agent_step`. If the on-duty agent decides to transfer control to another agent, the `handoff` method defines `next_agent` in Context and uses `DEFAULT_HANDOFF_OUTPUT_PROMPT` to inform the succeeding agent to continue handling the user request. ![If an agent finds that it can't handle the user's request, it will use the handoff method to transfer control.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.handoff.drawio.png) If an agent finds that it can't handle the user's request, it will use the handoff method to transfer control. Image by Author `parse_agent_output` interprets executable tools; if none remain, the workflow returns the final result. Otherwise, it initiates concurrent execution. `call_tool` finds and executes the specific tool’s code, writing results into `ToolCallResult` and throwing a copy into the output stream. `aggregate_tool_results` consolidates tool call results, and checks if `handoff` was executed – if so, switch to the next on-duty agent, restarting the process. Otherwise, if no `handoff` or the tool's `return_redirect` is False, it restarts. Other scenarios end Workflow, while calling agent's `handle_tool_call_results` and `finalize` allows adjusting language model outcomes. Apart from standard Workflow step methods, AgentWorkflow includes a `from_tools_or_functions` method for easy name comprehension. When using AgentWorkflow as an independent Agent, this initiates calling FunctionAgent or ReActAgent, executing them. Here’s an example: ```python from tavily import AsyncTavilyClient async def search_web(query: str) -> str: """Useful for using the web to answer questions""" client = AsyncTavilyClient() return str(await client.search(query)) workflow = AgentWorkflow.from_tools_or_functions( [search_web], system_prompt="You are a helpful assistant that can search the web for information." ) ``` ### Useful Events in the Event Stream You might have noticed that after adopting a multi-agent orchestration framework, one of the biggest hurdles we face is the long wait time for the workflow to complete all agent executions, and it's hard to know what's happening during the workflow execution. The handoff mechanism of AgentWorkflow handles this much better: when an agent gains control, it continuously responds to user requests without having to re-execute the workflow each time. For visualizing the steps during workflow execution, AgentWorkflow solves this by throwing events in the stream output pipeline in real time. Similar to LlamaIndex Workflow, after calling the workflow's `run` method, we can use the `handler.stream_events()` method to get all the events in the pipeline, and then use the `isinstance` method to filter the events: ```python handler = workflow.run( user_msg=message.content, ctx=context ) stream_msg = cl.Message(content="") async for event in handler.stream_events(): if isinstance(event, AgentInput): print(f"========{event.current_agent_name}:=========>") print(event.input) print("=================<") if isinstance(event, AgentOutput) and event.response.content: print("<================>") print(f"{event.current_agent_name}: {event.response.content}") print("<================>") if isinstance(event, AgentStream): await stream_msg.stream_token(event.delta) await stream_msg.send() ``` Specifically, in the order of calls, AgentWorkflow throws five events: `AgentInput`, `AgentStream`, `AgentOutput`, `ToolCall`, and `ToolCallResult`, as shown in the diagram below: ![The yellow oval represents the events in stream_events.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.streaming_events.drawio-1.png) The yellow oval represents the events in stream\_events. Image by Author `AgentInput` is thrown in the `take_step` method of `FunctionAgent`, mainly containing the current chat history and agent name. Since the chat history is quite long, we only use this event for debugging and do not display it on the interface. For me, `AgentStream` is the most useful event because it outputs the intermediate results of the current agent call as a message stream. If you want to understand what the large language model is thinking during workflow execution, you can focus on this event. But this event also outputs many intermediate results you might not need, depending on your choice. ![The effect of streaming intermediate processes in AgentStream.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/agent_stream.gif) The effect of streaming intermediate processes in AgentStream. Image by Author `AgentOutput` is thrown by AgentWorkflow after the `take_step` method of `FunctionAgent` is completed. The main difference between this event and `AgentStream` is that it is a synchronous event. If you need to get all the messages of the current round at once, you can focus on this event. `ToolCall` and `ToolCallResult` are used to contain the parameters of the tool call and the results from the tool call side, respectively. Like `AgentInput`, since the messages in these two events are quite long, we only use them for debugging rather than displaying them on the interface. Having covered AgentWorkflow’s basics, we'll now move on to project practice. To offer a direct comparison, this project again uses the customer service example from previous articles, displaying how simple AgentWorkflow's development can be. --- ## Customer Service Project Practice Based on Agentworkflow In a previous article, I demonstrated using a customer service project to showcase LlamaIndex Workflow’s capability of multi-agent orchestration akin to OpenAI Swarm. Today's project uses AgentWorkflow to present its development ease with the same customer service project for clear understanding. ### Final effect Here’s the final project display: ![The final effect of this project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-1.png) The final effect of this project. Image by Author As shown, when a user makes a request, the system automatically hands it off to the corresponding agent based on intent. Next are the core codes. Due to length, only important code is presented here; visit the code repository at the article's end for details. ### Defining agents In the multi-agent-customer-service project, I’ll create a new `src_v2` folder and modify the `sys.path` in `app.py` to reuse the previously created data model. In the previous project, the customer demand response logic was written into Workflow, making `workflow.py` unwieldy and tough to maintain. This time, `ConciergeAgent`, `PreSalesAgent`, and `PostSalesAgent` will truly handle customer services, using AgentWorkflow framework code without business logic addition. Hence, a new `agents.py` file defines `concierge_agent`, `pre_sales_agent`, and `post_sales_agent` agent instances. ![We will define three separate agents.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.agents.drawio.png) We will define three separate agents. Image by Author Each agent requires a `name` and `description`, crucial as AgentWorkflow organizes them by these as key-value pairs for `handoff` references, determining the next agent transition. Starting with `concierge_agent`, it checks if the user has registered a name – if not, it executes the `login` tool for registration; otherwise, based on intent, it decides whether to transfer control to the other two agents. ```python concierge_agent = FunctionAgent( name="ConciergeAgent", description="An agent to register user information, used to check if the user has already registered their title.", system_prompt=( "You are an assistant responsible for recording user information." "You check from the state whether the user has provided their title or not." "If they haven't, you should ask the user to provide it." "You cannot make up the user's title." "If the user has already provided their information, you should use the login tool to record this information." ), tools=[login], can_handoff_to=["PreSalesAgent", "PostSalesAgent"] ) ``` Then comes `pre_sales_agent`, responsible for pre-sales inquiries. Upon receiving a request, it reviews chat history, queries `VectorIndex` according to inquiries, and responds strictly following documentation. If the user isn’t inquiring about pre-sales, it transfers control to the other two agents. ```python pre_sales_agent = FunctionAgent( name="PreSalesAgent", description="A pre-sales assistant helps answer customer questions about products and assists them in making purchasing decisions.", system_prompt=( "You are an assistant designed to answer users' questions about product information to help them make the right decision before purchasing." "You must use the query_sku_info tool to get the necessary information to answer the user and cannot make up information that doesn't exist." "If the user is not asking pre-purchase questions, you should transfer control to the ConciergeAgent or PostSalesAgent." ), tools=[query_sku_info], can_handoff_to=["ConciergeAgent", "PostSalesAgent"] ) ``` Lastly, `post_sales_agent` handles questions and after-sales policies regarding product usage. Like `pre_sales_agent`, it can only reply based on existing documents, minimizing large language model misconceptions. ```python post_sales_agent = FunctionAgent( name="PostSalesAgent", description="After-sales agent, used to answer user inquiries about product after-sales information, including product usage Q&A and after-sales policies.", system_prompt=( "You are an assistant responsible for answering users' questions about product after-sales information, including product usage Q&A and after-sales policies." "You must use the query_terms_info tool to get the necessary information to answer the user and cannot make up information that doesn't exist." "If the user is not asking after-sales or product usage-related questions, you should transfer control to the ConciergeAgent or PreSalesAgent." ), tools=[query_terms_info], can_handoff_to=["ConciergeAgent", "PreSalesAgent"] ) ``` ### Ui development with Chainlit Since Workflow logic is no longer necessary, after developing all agents, UI development can commence directly, again using Chainlit. In `ready_my_workflow`, initialize `AgentWorkflow` and `Context` while storing workflow and context instances in `user_session` in the start method: ```python def ready_my_workflow() -> tuple[AgentWorkflow, Context]: workflow = AgentWorkflow( agents=[concierge_agent, pre_sales_agent, post_sales_agent], root_agent=concierge_agent.name, initial_state={ "username": None } ) ctx = Context(workflow=workflow) return workflow, ctx @cl.on_chat_start async def start(): workflow, ctx = ready_my_workflow() cl.user_session.set("workflow", workflow) cl.user_session.set("context", ctx) await cl.Message( author="assistant", content=GREETINGS ).send() ``` Next, in the `main` method, fetch user messages and call workflow for responses. Additional code is provided to demonstrate monitoring `AgentInput` and `AgentOutput` message streams; adjust as needed: ```python @cl.on_message async def main(message: cl.Message): workflow: AgentWorkflow = cl.user_session.get("workflow") context: Context = cl.user_session.get("context") handler = workflow.run( user_msg=message.content, ctx=context ) stream_msg = cl.Message(content="") async for event in handler.stream_events(): if isinstance(event, AgentInput): print(f"========{event.current_agent_name}:=========>") print(event.input) print("=================<") if isinstance(event, AgentOutput) and event.response.content: print("<================>") print(f"{event.current_agent_name}: {event.response.content}") print("<================>") if isinstance(event, AgentStream): await stream_msg.stream_token(event.delta) await stream_msg.send() ``` With this, our project code is complete. AgentWorkflow encapsulates multi-agent orchestration logic well, making our v2 version more focused, where good agent writing suffices. --- Next, we will try to integrate LlamaIndex and DeepSeek-R1 to enable AgentWorkflow to output reasoning content: [Integrating LlamaIndex and DeepSeek-R1 for reasoning\_content and Function Call FeaturesEmpowering AgentWorkflow with the strong boost from DeepSeek-R1![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-24.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover.webp)](https://www.dataleadsfuture.com/integrating-llamaindex-and-deepseek-r1-for-reasoning%5Fcontent-and-function-call-features-2/) --- ## Improving FunctionAgent Executing my project code, you might notice something odd: ![The agent can't reply to the user's request in time and needs to ask twice.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-2.png) The agent can't reply to the user's request in time and needs to ask twice. Image by Author The system correctly identifies user intent and hands it to the next agent, but the latter doesn't immediately respond, requiring the user to repeat. I've written a detailed article diving deep into the technical reasons behind this phenomenon and the ultimate solution. You can read it here: [Fixing the Agent Handoff Problem in LlamaIndex’s AgentWorkflow SystemThe problem: agents that won’t continue after handoff![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-26.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/p60Imf7mlg8tNL8kC9VdJSw8AsixrM.webp)](https://www.dataleadsfuture.com/fixing-the-agent-handoff-problem-in-llamaindexs-agentworkflow-system/) After a series of debugs, I located the problem: the agent taking over cannot well trace back the chat history to find the user's request. Thus, I attempted to extend FunctionAgent and modify some codes. After some tweaks, agents now respond promptly upon receiving the handoff, proving effective: ![The post-sales agent takes over the user's request and replies immediately. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/image-3.png) The post-sales agent takes over the user's request and replies immediately. Image by Author Let me explain the reason and how I handled it: In the original implementation of FunctionAgent, each tool\_call and its result are stored as separate messages in the `ChatMemory`call from the handoff method. For example, when the agent on duty realizes it needs to hand off the user's request to the next agent, a message related to the handoff method call is stored. Once the handoff is completed, another message indicating the completion is stored. The user's request is already stored in the `ChatMemory` before these two handoff-related messages are added. `ChatMemory` is a first-in, first-out queue, meaning the earliest messages are at the front. It also has a token limit, so messages exceeding this limit are removed from the queue. After the agent handoff, two situations can occur: 1. The user's request message is still in the queue but pushed to the front by the two handoff messages, causing the large language model to overlook the user's request. 2. Due to the token limit, the user's request message is removed from the queue, preventing the large language model from recalling the user's request from `ChatMemory`. In either case, the new agent can't perceive the user's previous request and waits for the user to resubmit it. ![The user's request got kicked out because of the token limit.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/03/AgentWorkflow.kickout.drawio.png) The user's request got kicked out because of the token limit. Image by Author Understanding the issue, the solution was simple: relocate the user’s previous request to the chat history’s end after the agent transition. The `take_step` in FunctionAgent is called when AgentWorkflow’s `run_agent_step` begins. Hence, inserting adjustment logic here is ideal: ```python class MyFunctionAgent(FunctionAgent): @override async def take_step( self, ctx: Context, llm_input: List[ChatMessage], tools: Sequence[AsyncBaseTool], memory: BaseMemory, ) -> AgentOutput: last_msg = llm_input[-1] and llm_input[-1].content state = await ctx.get("state", None) print(f">>>>>>>>>>>{state}") if "handoff_result" in last_msg: for message in llm_input[::-1]: if message.role == MessageRole.USER: last_user_msg = message llm_input.append(last_user_msg) break return await super().take_step(ctx, llm_input, tools, memory) ``` As shown, I iterate `chat_history` in reverse until finding the most recent user-requested message, appending it to `chat_history’s` end. Then I can replace my Agent from the FunctionAgent class with MyFunctionAgent, like this: ```python concierge_agent = MyFunctionAgent( name="ConciergeAgent", description="An agent to register user information, used to check if the user has already registered their title.", system_prompt=( "You are an assistant responsible for recording user information." "You check from the state whether the user has provided their title or not." "If they haven't, you should ask the user to provide it." "You cannot make up the user's title." "If the user has already provided their information, you should use the login tool to record this information." ), tools=[login], can_handoff_to=["PreSalesAgent", "PostSalesAgent"] ) ``` You can do the same thing with pre-sales-agent and post-sales-agent. A potential challenge arises: how to apply this solely during agent transition, bypassing regular steps? Earlier, we noted that AgentWorkflow returns `handoff_output_prompt` after handoff executes the `handoff` method. The succeeding agent's most recent message is this `handoff_output_prompt`. Hence, during AgentWorkflow initialization, I pass in a custom `handoff_output_prompt` similar to the default but tagged upfront with "handoff\_result": ```python def ready_my_workflow() -> tuple[AgentWorkflow, Context]: workflow = AgentWorkflow( agents=[concierge_agent, pre_sales_agent, post_sales_agent], root_agent=concierge_agent.name, handoff_output_prompt=( "handoff_result: Due to {reason}, the user's request has been passed to {to_agent}." "Please review the conversation history immediately and continue responding to the user's request." ), initial_state={ "username": None } ) ctx = Context(workflow=workflow) return workflow, ctx ``` Thus, in `take_step`, user message relocation only occurs when messages include the `handoff_result` tag, effectively resolving the issue. --- ## Conclusion In today's increasingly rich multi-agent orchestration scenarios, LlamaIndex has timely adjusted its positioning and launched the AgentWorkflow framework last month, greatly simplifying the development of agent orchestration based on LlamaIndex Workflow. In today’s article, I thoroughly explained AgentWorkflow’s principles and illustrated through the customer service project practice how development has simplified compared to only using Workflow. Though I believe AgentWorkflow brings LlamaIndex’s multi-agent solution close to perfection, the framework's recent release means specific scenarios still need refinement. I look forward to further practices to enhance it. Keep pushing forward, LlamaIndex! Thanks for reading. You’re welcome to comment on your perspective regarding LlamaIndex AgentWorkflow, and I’ll respond as soon as possible. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- next step Mr. Qian's Recommendation: Designing an event-driven LlamaIndex Workflow is one of the biggest steps you can take. From here, you should think about how to use LlamaIndex to build a production-ready, high-accuracy RAG system. We recommend checking out the [****RAG for Generative AI Applications specialization**](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) course brought to you by industry experts at IBM. You'll get hands-on experience with all the key tools and skills you need to build a RAG Workflow, so you can hit the ground running from day one. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [Start Learning for Free Now ](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) --- Here are the associated source codes from this article for free reading or modification: [GitHub - qtalen/multi-agent-customer-serviceContribute to qtalen/multi-agent-customer-service development by creating an account on GitHub.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/pinned-octocat-093da3e6fa40-1.svg)GitHubqtalen![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/multi-agent-customer-service-1)](https://github.com/qtalen/multi-agent-customer-service?ref=dataleadsfuture.com) ### Using LLamaIndex Workflow to Implement an Agent Handoff Feature Like OpenAI Swarm URL: https://www.dataleadsfuture.com/using-llamaindex-workflow-to-implement-an-agent-handoff-feature-like-openai-swarm/ Last updated: 2026-07-06T10:01:34.000Z Happy Lunar New Year, my friends! In the last article, I introduced the Workflow framework of LlamaIndex. [Deep Dive into LlamaIndex Workflow: Event-driven LLM architectureWhat I think about the progress and shortcomings after practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color-icon-1.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LLamaIndex_Workflow.webp)](https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/) Today, I will show you how to use LlamaIndex Workflow to implement a multi-agent orchestration feature similar to [OpenAI Swarm](https://github.com/openai/swarm?ref=dataleadsfuture.com), using a customer service chatbot project as an example. ## Introduction Remember the Swarm framework released by OpenAI not long ago? Its biggest feature is agents and handoffs. The agents are straightforward: they use a set of specific commands and tools to get tasks done. It's like putting a LLM function call into a neat package. And handoffs are different. They allow an agent to pass the work to another agent seamlessly based on the context of the current conversation, making agents work together without any hiccups. ### Why this is important Let's look at a diagram explaining the whole process of a ReactAgent. ![The ReactAgent needs at least three accesses to LLM to complete.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/llamaindex_workflow-ReAct.drawio-1.png) The ReactAgent needs at least three accesses to LLM to complete. Image by Author Just a simple agent call, like one, two, three, needs at least three accesses to LLM to complete. Traditional agent applications are like this, keeping conversation context and user state, and the agent call chain is usually fixed. For each user request, agents have to call LLM multiple times to check the state, and honestly, some calls are unnecessary. Here's an example: imagine we have an e-commerce website, and we need a customer service team to answer users' questions. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-agent-chain.drawio.png) In an agent chain, agents are invoked sequentially. Image by Author In a chain agent application, every question from a user goes to the front desk, and then the front desk asks for the pre-sales service. If they can't answer, the front desk asks for after-sales service, and then the front desk reorganizes the answers from the backend and replies to the customer. Isn't that silly? Look at all the unnecessary delays and call costs it causes! ### How Swarm does it Swarm uses a handoff approach that fits the real world better. Let me use that customer service example again: ![Agent handoff allows you to interact directly with the corresponding customer service. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-swarm.drawio.png) Agent handoff allows you to interact directly with the corresponding customer service. Image by Author Imagine a store called Swarm. When a customer asks the front desk a question, the front desk figures out what kind of question it is (pre-sale or after-sale) and passes the customer to the corresponding service. Then, the customer talks to that service directly. Sounds reasonable, right? So why don't we just use Swarm? ### Why not just use Swarm Because Swarm is still just an experimental framework. According to the official statement: > Swarm is currently an experimental sample framework intended to explore ergonomic interfaces for multi-agent systems. It is not intended to be used in production and therefore has no official support. (This also means we will not be reviewing PRs or issues!) So, we can't use Swarm directly in production systems. But what we need is the agent handoff capability, right? Since that's the case, why not build a similar framework yourself? Today's article is written for this purpose. We will develop a project using a customer service system as an example, which will use Workflow to implement agent orchestration and handoff capabilities. Let's get started. --- ## Project in Practice: A Customer Service Chatbot with Agent Handoff Capability This project is quite complex. To help you understand my implementation, I have put the entire project code at the end of the article. You can freely read and modify it without my permission. 💡 Want to know more about my work in AI applications or the field of data science? Feel free to [****Subscribe Now**](#/portal/signup), everything is free! ### Step one, set up an interactive interface Whether you use an agent or not, you always need to adjust your prompts and code logic. At this point, a what-you-see-is-what-you-get chat UI becomes very important. In this section, I'll use chainlit to quickly implement a super cool web-based chat window. [Chainlit](https://docs.chainlit.io/get-started/overview?ref=dataleadsfuture.com) is a Python library built on [Streamlit](https://streamlit.io/?ref=dataleadsfuture.com). This means you don't need any frontend skills to quickly build a Chatbot prototype. (Hooray) Let's get moving. ![The scaffold of our project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/image.png) The scaffold of our project. Image by Author First, we create a `.env` file in the project's root directory, which stores important environmental variables like `OPENAI_API_KEY` and `OPENAI_BASE_URL`. Later, I will use [dotenv](https://pypi.org/project/python-dotenv/?ref=dataleadsfuture.com) to read it. This is important because by using the `.env` file, you can strip the `API_KEY` from your code, then you can freely publish your code. Next, we need to set up a simple project scaffold. Our project will contain two folders: `src` and `data`. Our Python source code files will be placed in the `src` folder, while text source files for RAG use will be placed in the `data` folder. In the `src` directory, first create an `app.py` file, which will act as the view to launch the `chainlit` interface. This file consists of three parts: 1. Code to prepare the Workflow program. 2. Code to respond to the user lifecycle, outputting intermediate processes. 3. Actual code to call the Workflow agent and conduct the conversation. The code flowchart is shown below: ![Flowchart of the project UI interface.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-interface.drawio.png) Flowchart of the project UI interface. Image by Author As a production-ready system, we often need to connect to the enterprise-private deployment of large model ports. How to connect to a private large model can be referred to in this article. [How to Connect LlamaIndex with Private LLM API DeploymentsWhen your enterprise doesn’t use public models like OpenAI![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color-icon-1-1.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LlamaIndex_LLM_API.drawio.png)](https://www.dataleadsfuture.com/how-to-connect-llamaindex-with-private-llm-api-deployments/) To make our customer service less rigid, we can set the temperature a bit higher. Here is the code for initializing the system environment, I will talk about the implementation of `CustomerService` later: ```python llm = OpenAILike( model="qwen-max-latest", is_chat_model=True, is_function_calling_model=True, temperature=0.35 ) Settings.llm = llm ``` Imagine, when the next customer service takes over to answer your question, what will she do first? Right, she needs to check the conversation history first. So we need to create a unique, conversation-context and user-state-preserving workflow for each distinguished user in the user session: ```python GREETINGS = "Hello, what can I do for you?" def ready_my_workflow() -> CustomerService: memory = ChatMemoryBuffer( llm=llm, token_limit=5000 ) agent = CustomerService( memory=memory, timeout=None, user_state=initialize_user_state() ) return agent def initialize_user_state() -> dict[str, str | None]: return { "name": None } @cl.on_chat_start async def start(): workflow = ready_my_workflow() cl.user_session.set("workflow", workflow) await cl.Message( author="assistant", content=GREETINGS ).send() ``` At the same time, I will also use chainlit's `cl.step` decorator to implement a simple logging method, which can help us output some process logs on the page, letting users know where we are now: ```python @cl.step(type="run", show_input=False) async def on_progress(message: str): return message ``` Then there is the `main` method, which is called every round of conversation. ```python @cl.on_message async def main(message: cl.Message): workflow: CustomerService = cl.user_session.get("workflow") context = cl.user_session.get("context") msg = cl.Message(content="", author="assistant") user_msg = message.content handler = workflow.run( msg=user_msg, ctx=context ) async for event in handler.stream_events(): if isinstance(event, ProgressEvent): await on_progress(event.msg) await msg.send() result = await handler msg.content = result await msg.update() cl.user_session.set("context", handler.ctx) ``` In this method, we first get the user-inputted dialogue, then call the workflow's run method to start the agent routing, while iterating through the events in the workflow pipeline and calling `on_progress` to output to the page. Finally, we output the result of the dialogue on the page and update the Context. To match the construction of the chainlit interface, we can first write a simple workflow: ```python class CustomerService(Workflow): def __init__( self, llm: OpenAILike | None = None, memory: ChatMemoryBuffer = None, user_state: dict[str, str | None] = None, *args, **kwargs ): self.llm = llm or Settings.llm self.memory = memory or ChatMemoryBuffer() self.user_state = user_state super().__init__(*args, **kwargs) @step async def start(self, ctx: Context, ev: StartEvent) -> StopEvent: ctx.write_event_to_stream(ProgressEvent(msg="We're making some progress.")) return StopEvent(result="Hello World") ``` Tada, our interactive interface is out: ![Our UI interface for this project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/image-1.png) Our UI interface for this project. Image by Author Next, we can start preparing the ingredients for today, and text source files for RAG use. ### Step two, generate text files Since this project is about simulating a customer support team for an online drone e-commerce website, I plan to set the background to an online unmanned aerial vehicle e-commerce site. I need two files: one file to introduce the drones being sold in the store and their details. Another file contains common FAQs about drone use and after-sales terms. To avoid business and data licensing issues, I plan to use LLM to generate the text I want. I specifically instructed LLM not to include any brands or real product information. Here is a screenshot of my file generation: ![Screenshot of data file generated using LLM.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/image-2.png) Screenshot of data file generated using LLM. Image by Author You can take my prompt as a reference: ```python SKUS_TEMPLATE_EN = """ You are the owner of an online drone store, please generate a description in English of all the drones for sale. Include the drone model number, selling price, detailed specifications, and a detailed description in more than 400 words. Do not include brand names. No less than 20 types of drones, ranging from consumer to industrial use. """ ``` ```python TERMS_TEMPLATE_EN = """ You are the head of a brand's back office department, and you are asked to generate a standardized response to after-sales FAQs in English that is greater than 25,000 words. The text should include common usage questions, as well as questions related to returns and repairs after the sale. This text will be used as a reference for the customer service team when answering customer questions about after-sales issues. Only the body text is generated, no preamble or explanation is added. """ ``` ### Step three, handle indexing and retrieve privatized data The foundation LLM does not contain corporate internal data. For enterprise applications, it is inevitable to use RAG to allow the LLM to access corporate privatized data. Our drone store is no exception. Before letting the agent staff start work, we need to provide them with some tools to access the product catalog and after-sales policy. LlamaIndex provides many indexes suitable for different occasions. If used in a real system, I would prefer to use `KnowledgeGraphIndex` for product information text. However, to make the sample project easy to understand, I still choose to use `chromadb` and `VectorStoreIndex`: ```python def get_index(collection_name: str, files: list[str]) -> VectorStoreIndex: chroma_client = chromadb.PersistentClient(path="temp/.chroma") collection = chroma_client.get_or_create_collection(collection_name) vector_store = ChromaVectorStore(chroma_collection=collection) storage_context = StorageContext.from_defaults(vector_store=vector_store) ready = collection.count() if ready > 0: print("File already loaded") index = VectorStoreIndex.from_vector_store(vector_store=vector_store) else: print("File not loaded.") docs = SimpleDirectoryReader(input_files=files).load_data() index = VectorStoreIndex.from_documents( docs, storage_context=storage_context, embed_model=embed_model, transformer=[SentenceSplitter(chunk_size=512, chunk_overlap=20)] ) return index INDEXES = { "SKUS": get_index("skus_docs", ["data/skus_en.txt"]), "TERMS": get_index("terms_docs", ["data/terms_en.txt"]) } ``` The running flowchart of this code is as follows: ![The running flowchart of the code.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-models.drawio.png) The running flowchart of the code. Image by Author If vector data already exists, return the index directly. If the data has not been loaded yet, first load the data into the vector store, then return the index. Then we add a tool method to help the agent get the corresponding retriever: ```python async def query_docs( index: VectorStoreIndex, query: str, similarity_top_k: int = 1 ) -> str: retriever = index.as_retriever(similarity_top_k=similarity_top_k) nodes = await retriever.aretrieve(query) result = "" for node in nodes: result += node.get_content() + "\n\n" return result ``` 💡 Want to know more about my work in AI applications or the field of data science? Feel free to [****Subscribe Now**](#/portal/signup), everything is free! ### Step four, hire a few agents Since we are building a smart customer service project, it is necessary to hire a few customer service agents. We also need to set up a data class for agents, which needs to include instructions and a set of tools, just like Swarm agents. Here we use `AgentConfig` to constrain, which is inherited from Pydantic's `BaseModel`. ```python class AgentConfig(BaseModel): """ Detailed configuration for an agent """ model_config = ConfigDict(arbitrary_types_allowed=True) name: str = Field(description="agent name") description: str = Field( description="agent description, which describes what the agent does" ) system_prompt: str | None = None tools: list[BaseTool] | None = Field( description="function tools available for this agent" ) ``` We need to hire a lobby manager agent, this agent needs to mark which agent the user will be handed over to. ```python class TransferToAgent(BaseModel): """Used to explain which agent to transfer to next.""" agent_name: str = Field(description="The name of the agent to transfer to.") ``` We also need to design a request transfer agent, which uses this agent to notify the workflow when a certain customer service cannot answer the user's question. ```python class RequestTransfer(BaseModel): """ Used to indicate that you don't have the necessary permission to complete the user's request, or that you've already completed the user's request and want to transfer to another agent. """ pass ``` Then we need to prepare a few tools for the agents to use: First is a `login` tool, this tool is only used to register the user's name. If you need to handle the user's login action, you can implement the details in this method. I use a closure to return a tool list. ```python def get_authentication_tools() -> list[BaseTool]: async def login(ctx: Context, username: str) -> bool: """When the user provides their name, you can use this method to update their status.。 :param username The user's title or name. """ if not username: return False user_state = await ctx.get("user_state", None) user_state["name"] = username.strip() await ctx.set("user_state", user_state) return True return [FunctionToolWithContext.from_defaults(async_fn=login)] ``` Here is a detail, since the tool needs to handle the user state saved in the workflow Context, we need to access the `ctx` object. But when the tool is called by the agent, the agent cannot sense the `ctx` object, so we need to let the agent ignore it. Here I modified the behavior of the `FunctionTool` module of LlamaIndex and rewrote the `FunctionToolWithContext` module. To save time, I referred to [an example](https://github.com/run-llama/multi-agent-concierge/blob/main/utils.py?ref=dataleadsfuture.com) on the official website, you can find it here. Of course, you can also find the source code at the end of the project code in the article. We also need tools to get the product catalog and after-sales terms, these two tools are direct calls to the retriever, quite simple. ```python def get_pre_sales_tools() -> list[BaseTool]: async def skus_info_retrieve(ctx: Context, query: str) -> str: """ When the user asks about a product, you can use this tool to look it up. :param query: The user's request. :return: The information found. """ sku_info = await query_docs(INDEXES["SKUS"], query) return sku_info return [FunctionToolWithContext.from_defaults(async_fn=skus_info_retrieve)] def get_after_sales_tools() -> list[BaseTool]: async def terms_info_retrieve(ctx: Context, query: str) -> str: """ When the user asks about how to use a product, or about after-sales and repair options, you can use this tool to look it up. :param query: The user's request. :return: The information found. """ terms_info = await query_docs(INDEXES["TERMS"], query) return terms_info return [FunctionToolWithContext.from_defaults(async_fn=terms_info_retrieve)] ``` To sum up, we need three professional customer service agents: 1. The first one is the front desk, used to register the user's visit. 2. The second one is the pre-sales service, used to recommend various products to users. 3. The third one is the after-sales service, used to answer various usage questions and after-sales terms. ```python def _get_agent_configs() -> list[AgentConfig]: return [ AgentConfig( name="Authentication Agent", description="Record the user's name. If there's no name, you need to ask this from the customer.", system_prompt=""" You are a front desk customer service agent for registration. If the user hasn't provided their name, you need to ask them. When the user has other requests, transfer the user's request. """, tools=get_authentication_tools() ), AgentConfig( name="Pre Sales Agent", description="When the user asks about product information, you need to consult this customer service agent.", system_prompt=""" You are a customer service agent answering pre-sales questions for customers. You will respond to users' inquiries based on the context of the conversation. When the context is not enough, you will use tools to supplement the information. You can only handle user inquiries related to product pre-sales. Please use the RequestTransfer tool to transfer other user requests. """, tools=get_pre_sales_tools() ), AgentConfig( name="After Sales Agent", description="When the user asks about after-sales information, you need to consult this customer service agent.", system_prompt=""" You are a customer service agent answering after-sales questions for customers, including how to use the product, return and exchange policies, and repair solutions. You respond to users' inquiries based on the context of the conversation. When the context is not enough, you will use tools to supplement the information. You can only handle user inquiries related to product after-sales. Please use the RequestTransfer tool to transfer other user requests. """, tools=get_after_sales_tools() ) ] ``` According to the needs of the workflow, we also need to write two methods to register agents to the workflow. ```python def get_agent_config_pair() -> dict[str, AgentConfig]: agent_configs = _get_agent_configs() return {agent.name: agent for agent in agent_configs} def get_agent_configs_str() -> str: agent_configs = _get_agent_configs() pair_list = [f"{agent.name}: {agent.description}" for agent in agent_configs] return "\n".join(pair_list) ``` Finally, we also need a workflow `system_prompt` for orchestration, which will contain all agent information and user status, and hand over the user to the correct customer service agent when needed. This prompt will be used directly by the workflow, no separate agent is needed, so just put the prompt here: ```python ORCHESTRATION_PROMPT = """ You are a customer service manager for a drone store. Based on the user's current status, latest request, and the available customer service agents, you help the user decide which agent to consult next. You don't focus on the dependencies between agents; the agents will handle those themselves. If the user asks about something unrelated to drones, you should politely and briefly decline to answer. Here is the list of available customer service agents: {agent_configs_str} Here is the user's current status: {user_state_str} """ ``` ### Step five, build the core workflow After so much preparation, we can finally get to the main course, and I'm sure everyone is eager to get started haha. Since the workflow is an event-driven framework, we need to define several events like before: ```python class OrchestrationEvent(Event): query: str class ActiveSpeakerEvent(Event): query: str class ToolCallEvent(Event): tool_call: ToolSelection tools: list[BaseTool] class ToolCallResultEvent(Event): chat_message: ChatMessage class ProgressEvent(Event): msg: str ``` 1. `OrchestrationEvent` to indicate that the workflow needs to transfer agents. 2. After the agent transfer is completed, `ActiveSpeakerEvent` will tell the workflow to use the new agent to answer the user. 3. If the agent needs to make a Function call, `ToolCallEvent` will be thrown to execute concurrently. 4. The results of concurrent execution will be thrown out with `ToolCallResultEvent`, and summarized into the final result. 5. Finally, we also need a `ProgressEvent` to stream the intermediate steps, making it easy for users to know where we are now. To avoid too much information interference, we only output the information of agent transfer here. After defining various events, we need to start writing the workflow. The workflow this time is a bit complex, so to make it easier for everyone to understand, I still drew a flowchart: ![The flowchart of our workflow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-workflow.drawio.png) The flowchart of our workflow. Image by Author First, let's look at the `start` method. The `start` method is relatively simple, as the entry method for user dialogue, it is responsible for storing the user's message in `ChatMemory`, and then judging whether there is an available agent currently, if there is, it throws an `ActiveSpeakerEvent`, entering the next step, if not, it throws an `OrchestrationEvent`, entering the agent orchestration. ![The flowchart of the start method.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-start.drawio.png) The flowchart of the start method. Image by Author ```python class CustomerService(Workflow): ... @step async def start( self, ctx: Context, ev: StartEvent ) -> ActiveSpeakerEvent | OrchestrationEvent: self.memory.put(ChatMessage( role="user", content=ev.msg )) user_state = await ctx.get("user_state", None) if not user_state: await ctx.set("user_state", self.user_state) user_msg = ev.msg active_speaker = await ctx.get("active_speaker", default=None) if active_speaker: return ActiveSpeakerEvent(query=user_msg) else: return OrchestrationEvent(query=user_msg) ``` From shallow to deep, let's look at how the agent orchestration `orchestrate` method works: ![The flowchart of the orchestrate method.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-orchestrate.drawio.png) The flowchart of the orchestrate method. Image by Author This method first takes the currently available agent and user state, and updates our `ORCHESTRATION_PROMPT` written in `agents.py`, thus getting the complete `system_prompt`. Then we use `TransferToAgent`, `system_prompt`, and `chat_history` all to LLM, letting LLM judge which agent to hand over to next. ```python class CustomerService(Workflow): ... @step async def orchestrate( self, ctx: Context, ev: OrchestrationEvent ) -> ActiveSpeakerEvent | StopEvent: chat_history = self.memory.get() user_state_str = await self._get_user_state_str(ctx) system_prompt = ORCHESTRATION_PROMPT.format( agent_configs_str=get_agent_configs_str(), user_state_str=user_state_str ) messages = [ChatMessage(role="system", content=system_prompt)] + chat_history tools = [get_function_tool(TransferToAgent)] event, tool_calls, _ = await self.achat_to_tool_calls(ctx, tools, messages) if event is not None: return event tool_call = tool_calls[0] selected_agent = tool_call.tool_kwargs["agent_name"] await ctx.set("active_speaker", selected_agent) ctx.write_event_to_stream( ProgressEvent(msg=f"In step orchestrate:\nTransfer to agent: {selected_agent}") ) return ActiveSpeakerEvent(query=ev.query) ``` After getting the latest agent, we update the context and throw an `ActiveSpeakerEvent`. We also need to define `achat_to_tool_calls` and `_get_user_state_str` as these two tool methods. `achat_to_tool_calls` method is responsible for getting the currently needed tools from LLM. `_get_user_state_str` is used to convert the user state into a string. ```python class CustomerService(Workflow): ... async def achat_to_tool_calls(self, ctx: Context, tools: list[FunctionTool], chat_history: list[ChatMessage] ) -> tuple[StopEvent | None, list[ToolSelection], ChatResponse]: response = await self.llm.achat_with_tools(tools, chat_history=chat_history) tool_calls: list[ToolSelection] = self.llm.get_tool_calls_from_response( response=response, error_on_no_tool_call=False ) stop_event = None if len(tool_calls) == 0: await self.memory.aput(response.message) stop_event = StopEvent( result=response.message.content ) return stop_event, tool_calls, response @staticmethod async def _get_user_state_str(ctx: Context) -> str: user_state = await ctx.get("user_state", None) user_state_list = [f"{k}: {v}" for k, v in user_state.items()] return "\n".join(user_state_list) ``` After studying the `Orchestrate` branch, let's see how the `ActiveSpeaker` branch works, which is the `speak_with_sub_agent` method: ![The flowchart of the speak_with_sub_agent method.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/multi-agent-customer-service-sub-agent.drawio.png) The flowchart of the speak\_with\_sub\_agent method. Image by Author This method first gets the current service-providing agent, as well as `chat_history` and `user_state`. Then use the current agent's `sys_prompt` and `tools`, as well as `chat_history`, to let LLM judge which tools to call next. ```python class CustomerService(Workflow): ... @step async def speak_with_sub_agent( self, ctx: Context, ev: ActiveSpeakerEvent ) -> OrchestrationEvent | ToolCallEvent | StopEvent: active_speaker = await ctx.get("active_speaker", default="") agent_config: AgentConfig = get_agent_config_pair()[active_speaker] chat_history = self.memory.get() user_state_str = await self._get_user_state_str(ctx) system_prompt = ( agent_config.system_prompt.strip() + f"\n\n:\n{user_state_str}" ) llm_input = [ChatMessage(role="system", content=system_prompt)] + chat_history tools = [get_function_tool(RequestTransfer)] + agent_config.tools event, tool_calls, response = await self.achat_to_tool_calls(ctx, tools, llm_input) if event is not None: return event await ctx.set("num_tool_calls", len(tool_calls)) for tool_call in tool_calls: if tool_call.tool_name == "RequestTransfer": await ctx.set("active_speaker", None) ctx.write_event_to_stream( ProgressEvent(msg="The agent is requesting a transfer, please hold on...") ) return OrchestrationEvent(query=ev.query) else: ctx.send_event( ToolCallEvent(tool_call=tool_call, tools=agent_config.tools) ) await self.memory.aput(response.message) ``` It is important to note that although the sample project is relatively simple, each agent's tools only have one, but in actual projects, there are often multiple tools to be called concurrently. So we need to iterate through `tool_calls`, throwing `ToolCallEvent` separately. At the same time, we also need to consider the situation where the current agent may not be able to handle the user's request, thus calling `RequestTransfer`, then we need to go back to the `orchestrate` step and re-select the agent. Let's look at the `handle_tool_calls` section, the code in this section looks a lot, but the actual thing to do is very simple, just get the tool to be executed and execute it, so simple that I don't even want to draw a flowchart. ```python class CustomerService(Workflow): ... @step(num_workers=4) async def handle_tool_calls( self, ctx: Context, ev: ToolCallEvent ) -> ToolCallResultEvent: tool_call = ev.tool_call tools_by_name = {tool.metadata.get_name(): tool for tool in ev.tools} tool_msg = None tool = tools_by_name[tool_call.tool_name] additional_kwargs = { "tool_call_id": tool_call.tool_id, "name": tool.metadata.get_name() } if not tool: tool_msg = ChatMessage( role="tool", content=f"Tool {tool_call.tool_name} does not exists.", additional_kwargs=additional_kwargs ) return ToolCallResultEvent(chat_message=tool_msg) try: if isinstance(tool, FunctionToolWithContext): tool_output = await tool.acall(ctx, **tool_call.tool_kwargs) else: tool_output = await tool.acall(**tool_call.tool_kwargs) tool_msg = ChatMessage( role="tool", content=tool_output.content, additional_kwargs=additional_kwargs ) except Exception as e: tool_msg = ChatMessage( role="tool", content=f"Encountered error in tool call: {e}", additional_kwargs=additional_kwargs ) return ToolCallResultEvent(chat_message=tool_msg) ``` There is a small detail here, I set a parameter `num_workers=4` for the `step` decorator. This is to tell the workflow that the concurrency is only up to 4, to avoid too high concurrency causing downstream system blockage. Then we come to the last method `aggregate_too_results`. ```python class CustomerService(Workflow): ... @step async def aggregate_tool_results( self, ctx: Context, ev: ToolCallResultEvent ) -> ActiveSpeakerEvent | None: num_tool_calls = await ctx.get("num_tool_calls") results = ctx.collect_events(ev, [ToolCallResultEvent] * num_tool_calls) if not results: return None for result in results: await self.memory.aput(result.chat_message) return ActiveSpeakerEvent(query="") ``` This method is relatively simple, just get all the execution results of `tool_calls`, then write them into `ChatMemory`, and finally hand them back to the agent to evaluate the results and answer the user. ### Step six, check our hard work At this point in the plot, our customer service team has been built, let's check if these agents are working hard. Start! ```bash chainlit run src/app.py ``` Not bad, the front desk agent first asks for my name, simulating the login process, then based on my needs, hands it over to the pre-sales agent: ![When my request was handed off to the pre-sales agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/image-3.png) When my request was handed off to the pre-sales agent. Image by Author It can also be transferred to the after-sales agent based on my request: ![When my request was handed off to the aftermarket agent.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2025/01/image-4.png) When my request was handed off to the aftermarket agent. Image by Author I can even play with them all day! --- What's next? The multi-agent orchestration of LlamaIndex AgentWorkflow will bring a more perfect solution: [Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration SolutionFurther optimization is introduced in this article![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-2.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4-1.webp)](https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/) --- ## Conclusion To summarize today's lesson, in this customer service chatbot project, we successfully simulated the agent handoff capability of OpenAI Swarm using LlamaIndex Workflow, achieving seamless collaboration among multiple agents. We can see that the agents handoff brings obvious advantages to the project: 1. The workflow autonomously decides which agent to use based on the dialogue context, there is no fixed code process, and it is completely decided by LLM. 2. Once the workflow decides which agent will serve the user, the user interacts directly with the corresponding agent, without any intermediate steps. However, there are still some shortcomings in the project: 1. The call to the agent is too low-level, causing us to handle the function calling process in the code. 2. The modularity of the workflow is not well done, I have also mentioned this in previous articles, which brings some obstacles to team collaboration. I look forward to gradually solving these problems in the upcoming articles. I welcome your comments and discussions and will reply to everyone as soon as possible. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- next step Mr. Qian's Recommendation: Designing an event-driven LlamaIndex Workflow is one of the biggest steps you can take. From here, you should think about how to use LlamaIndex to build a production-ready, high-accuracy RAG system. We recommend checking out the [****RAG for Generative AI Applications specialization**](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) course brought to you by industry experts at IBM. You'll get hands-on experience with all the key tools and skills you need to build a RAG Workflow, so you can hit the ground running from day one. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [Start Learning for Free Now ](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) ### Deep Dive into LlamaIndex Workflow: Event-driven LLM architecture URL: https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/ Last updated: 2026-07-06T09:56:30.000Z Recently, LlamaIndex introduced a new feature called [Workflow](https://docs.llamaindex.ai/en/stable/understanding/workflows/?ref=dataleadsfuture.com) in one of its versions, providing event-driven and logic decoupling capabilities for LLM applications. In today's article, we'll take a deep dive into this feature through a practical mini-project, exploring what's new and still lacking. Let's get started. --- ## Introduction ### Why event-driven? More and more LLM applications are shifting towards intelligent agent architectures, expecting LLMs to meet user requests through calling different APIs or multiple iterative calls. This shift, however, brings a problem: as agent applications make more API calls, program responses slow down and code logic becomes more complex. A typical example is [ReActAgent](https://docs.llamaindex.ai/en/stable/api%5Freference/agent/react/?ref=dataleadsfuture.com#llama%5Findex.core.agent.react.ReActAgent), which involves steps like Thought, Action, Observation, and Final Answer, requiring at least three LLM calls and one tool call. If loops are needed, there will be even more I/O calls. ![A typical ReAct agent will make at least three calls to LLM.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-ReAct.drawio-1.png) A typical ReAct agent will make at least three calls to LLM. Image by Author ### Is there a way to optimize this? As shown in the diagram above, in a traditional programming model, all I/O calls are linear; the next task must wait until the previous one is completed. Although mainstream LLMs now support result generation via stream output, in agent applications, we still need to wait for the LLM to finish generating results before returning or moving to the next phase. Actually, we don’t need all I/O calls to proceed sequentially; they can be executed concurrently, as shown in the diagram below: ![In concurrent programming, multiple steps are executed in parallel.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-Concurrence.drawio-1.png) In concurrent programming, multiple steps are executed in parallel. Image by Author Does this diagram look familiar? Yes, Python's `asyncio` package provides the ability to execute I/O-bound tasks concurrently, and nearly all I/O-based APIs, including LLM clients, support concurrent execution. LlamaIndex's Workflow also utilizes the principles of concurrent programming. It goes further by not only encapsulating the details of the `asyncio` library but also providing an event mechanism that allows us to decouple different segments of the business process. Now that we understand the background, let's step through LlamaIndex Workflow with an actual project. --- ## First Impressions Before the main course, let's have an appetizer by familiarizing ourselves with the elements and basic principles through a simple code example. ### Importing necessary packages First, we need to import the necessary tools. Workflow is already included in the latest version of LlamaIndex, no separate installation is needed. ```python from llama_index.core.workflow import ( Event, StartEvent, StopEvent, Workflow, Context, step, ) ``` ### Defining some events Since Workflow is an event-driven framework, we should start by defining some events. To avoid inconsistencies, we can first define a `BaseEvent`, ensuring all events use the key `payload` for message passing. ```python class BaseEvent(Event): payload: str | dict | None ``` Let’s define our first event of the day: `SecondStepEvent` ```python class SecondStepEvent(BaseEvent): ... ``` ### Starting simple Next, let’s start coding our first Workflow program, which is a subclass of `Workflow` containing two methods: ```python class SimpleWorkflow(Workflow): @step async def start(self, ev: StartEvent) -> SecondStepEvent: return SecondStepEvent(payload=ev.payload) @step async def second_step(self, ev: SecondStepEvent) -> StopEvent: return StopEvent(result=ev.payload) ``` 1. The method `start` accepts a `StartEvent` and then returns a `SecondStepEvent`. 2. The method `second_step` accepts a `SecondStepEvent` and then returns a `StopEvent`. Let's get the code up and running to see how it works. ```python s_wf = SimpleWorkflow(timeout=10, verbose=True) result = await s_wf.run(payload="hello world") print(result) ``` We have turned on the `verbose` option so that we can see in detail how the code is executed. ![The result of the execution of our first Workflow program.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image.png) The result of the execution of our first Workflow program. Image by Author ### Trying out the visualization tool LlamaIndex also generously provides a small tool that allows us to see the entire workflow process, which is very intuitive. ```python from llama_index.utils.workflow import draw_all_possible_flows draw_all_possible_flows(SimpleWorkflow, filename="simple_workflow.html") ``` ![Flowchart of the first Workflow code.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-1.png) Flowchart of the first Workflow code. Image by Author ### Explaining the principles A quick look at the source code reveals that Workflow internally maintains a `Context`, which not only keeps an event queue but also maintains a dictionary containing each step. ![Workflow uses a run_flow loop to listen for events and execute steps.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-Workflow-Cycle.drawio.png) Workflow uses a run\_flow loop to listen for events and execute steps. Image by Author When Workflow is initialized, the `step` decorator analyzes the signature of each method to determine which events it will receive and return, starting to listen to the event queue, and then storing this method in the `step` dictionary. When the Workflow's `run` method is launched, it starts a `runflow` loop, initially placing a `StartEvent` in the event queue. If there's a method that accepts this `StartEvent`, it starts executing and returns the corresponding event, putting it back into the event queue. The `step` method can also directly call the Context's `send_event` method to place an event in the queue. If the runflow loop detects a `StopEvent` in the queue, it exits the flow and returns the final result. With a basic understanding of the elements and implementation principles, we can now explore the advantages and shortcomings of the Workflow through a hands-on project. 💡 Want to know more about my work in AI applications or the field of data science? Feel free to [****Subscribe Now**](#/portal/signup), everything is free! --- ## Hands-on Project In today's hands-on project, we'll step by step implement an automated trading robot that listens to market sentiment and executes stock trades, demonstrating Workflow's branching and looping control, Streaming events, and concurrent execution features. *Note: The following code uses pseudocode to explain the Workflow mechanism and does not contain any trading logic or investment advice.* --- If you're interested in learning more about the complete LlamaIndex workflow, check out this tutorial I wrote: [Use LLamaIndex Workflow to Create an Ink Painting Style Image Generation WorkflowAdd strong artistic flair through fine control of LLM context![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-37.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/generated_00.webp)](https://www.dataleadsfuture.com/use-llamaindex-workflow-to-create-an-ink-painting-style-image-generation-workflow/) --- ### Branching and looping control In the first version of the trading robot, we'll continuously monitor the latest news of a certain stock, analyze the sentiment implied in the news, and then make corresponding trades. The entire code logic is shown in the diagram below: ![Flowchart of our TradeMonitor program.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-TradeMonitor.drawio.png) Flowchart of our TradeMonitor program. Image by Author First, we'll define a `Trading` class that uses `async` to implement the `buy` and `sell` methods. ```python class StockTrader: async def buy(self, stock: str) -> None: await asyncio.sleep(0.5) print(f"Buying {stock}") async def sell(self, stock: str) -> None: await asyncio.sleep(0.5) print(f"Selling {stock}") ``` We also need to implement four events: `LoopEvent`, `GetSentimentEvent`, `BuyEvent`, and `SellEvent`, all of which are subclasses of `BaseEvent`, ensuring they follow a unified message-passing interface. ```python class LoopEvent(BaseEvent): ... class GetSentimentEvent(BaseEvent): ... class BuyEvent(BaseEvent): ... class SellEvent(BaseEvent): ... ``` Next, we start implementing the `TradeMonitorWorkflow` class, which contains the core business logic. ```python class TradeMonitorWorkflow(Workflow): def __init__(self, total_cycle: int = 1, *args, **kwargs) -> None: self.total_cycle = total_cycle self.counter = 0 self.trader = StockTrader() super().__init__(*args, **kwargs) @step async def begin(self, ev: StartEvent | LoopEvent) \ -> GetSentimentEvent | StopEvent: print("We now return to the begin step") if isinstance(ev, StartEvent): self.stock = ev.payload if self.counter < self.total_cycle: await asyncio.sleep(3) self.counter += 1 return GetSentimentEvent(payload=self.stock) else: return StopEvent(result="We're done for the day.") @step async def get_sentiment(self, ev: GetSentimentEvent) -> BuyEvent | SellEvent: print(f"Wil get the latest sentiment for stock {ev.payload}") if random.random() < 0.3: return SellEvent(payload='Bearish') else: return BuyEvent(payload='Bullish') @step async def buy(self, ev: BuyEvent) -> LoopEvent: print(f"We now buy some stock with sentiment {ev.payload}.") await self.trader.buy(self.stock) return LoopEvent(payload="Start a new cycle.") @step async def sell(self, ev: SellEvent) -> LoopEvent: print(f"We now sell some stock with sentiment {ev.payload}") await self.trader.sell(self.stock) return LoopEvent(payload="Start a new cycle.") ``` 1. The `begin` method is our entry point, accepting `StartEvent` and `LoopEvent`. 2. The `StartEvent` is the default event that starts the code, and we pass the stock code through this event. 3. The `GetSentimentEvent` triggers the `get_sentiment` method to obtain sentiment information. For simplicity, we use the `random` method to generate two sentiments, `Bullish` and `Bearish`, and then return the corresponding `BuyEvent` or `SellEvent` based on the sentiment. 4. After a transaction is completed, the `LoopEvent` reinitiates the `begin` method for a new round of looping. To simplify the code, we set only one loop. 5. In each loop, the `begin` method returns a `GetSentimentEvent` to trigger the acquisition of the latest stock sentiment. If all loops are completed, it returns a `StopEvent`. 6. When a `BuyEvent` or `SellEvent` is received, the corresponding `step` method executes the transaction based on the sentiment flag in the message body and returns a `LoopEvent` to start a new loop. As you can see, by using events, we can decouple complex loops and branching processes, making it possible for corresponding events to trigger new loops. Let's use the `draw_all_possible_flows` tool to see if the entire flow chart matches our designed business logic diagram. ```python draw_all_possible_flows(TradeMonitorWorkflow, filename="trade_monitor_workflow.html") ``` ![We use events to decouple branching and looping control.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-2.png) We use events to decouple branching and looping control. Image by Author Is that all? If it's just about decoupling loops and branching processes, couldn't I achieve that with some coding tricks? Yes, but flow control is just the most superficial layer. Next, let's experience the powerful potential unleashed by combining `asyncio` with Workflow. ### Streaming events When building an agent chain, one of the most headache-inducing issues is how to feed back messages during the execution process to users, helping them understand the progress of code execution. In the code above, we use the `print` method to print progress in real-time on the console, but this approach is not feasible for a web applications. One solution is to launch a separate pipeline to push messages to users in real-time, but when multiple steps are executed concurrently, how to handle this pipeline becomes a challenge. Fortunately, the Workflow's Context directly provides a message streaming pipeline, and we can conveniently write messages into this pipeline and handle them uniformly at the calling end through an `async for` loop. ![LlamaIndex Workflow uses a streaming queue to output messages.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-Streaming-Event.drawio.png) LlamaIndex Workflow uses a streaming queue to output messages. Image by Author Let's modify our previous trading program: ```python class ProgressEvent(BaseEvent): ... class TradeMonitorWorkflowV2(Workflow): def __init__(self, total_cycle: int = 1, *args, **kwargs) -> None: self.total_cycle = total_cycle self.counter = 0 self.trader = StockTrader() super().__init__(*args, **kwargs) @step async def begin(self, ctx: Context, ev: StartEvent | LoopEvent) \ -> GetSentimentEvent | StopEvent: ctx.write_event_to_stream( ProgressEvent(payload="We now return to the begin step") ) ... @step async def get_sentiment(self, ctx: Context, ev: GetSentimentEvent) -> BuyEvent | SellEvent: ctx.write_event_to_stream( ProgressEvent(payload=f"Wil get the latest sentiment for stock {ev.payload}") ) ... @step async def buy(self, ctx: Context, ev: BuyEvent) -> LoopEvent: ctx.write_event_to_stream( ProgressEvent(payload=f"We now buy some stock with sentiment {ev.payload}.") ) ... @step async def sell(self, ctx: Context, ev: SellEvent) -> LoopEvent: ctx.write_event_to_stream( ProgressEvent(payload=f"We now sell some stock with sentiment {ev.payload}") ) ... ``` In the first step, we pass a `Context` type parameter in the signature of the `step` method. This lets Workflow know to pass the current execution context into the `step` method. Then, we replace the `print` method with the `ctx.write_event_to_stream` method to write messages into the pipeline in real-time. Finally, before waiting for the final result, we use the `stream_events` method to iterate over the latest messages from the message pipeline. ```python from datetime import datetime def streaming_log(message: str) -> None: current_time = datetime.now().strftime("%H:%M:%S") print(f"{current_time} {message}") trade_monitor_v2 = TradeMonitorWorkflowV2(timeout=10, verbose=False) handler = trade_monitor_v2.run(payload="[Stock Code]") async for event in handler.stream_events(): if isinstance(event , ProgressEvent): streaming_log(event.payload) final_result = await handler print("Final result: ", final_result) ``` ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-4.png) During code execution, Workflow streams out messages through the streaming queue. Image by Author ### Concurrent execution As mentioned at the beginning of the article, for I/O-bound tasks, we can use the `asyncio` package to make the code execute concurrently, greatly improving the running efficiency. Workflow implements this mechanism for us, encapsulating the `asyncio` execution code, and letting us focus on the code logic. Let's explain using the `TradeMonitor` project as an example. ![We can let multiple steps execute in parallel to optimize execution time.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-Concurrent-execution.drawio.webp) We can let multiple steps execute in parallel to optimize execution time. Image by Author This time, we'll upgrade the project, allowing the `TradeMonitor` to judge whether it's Bullish or Bearish not through one source but simultaneously through wallstreetjournal, wallstreetbets, and a machine learning trending predictor. First, we add six events: `WSBEvent`, `WSBSentimentEvent`, `WSJEvent`, `WSJSentimentEvent`, `TrendingPredictionEvent`, and `PredictionResultEvent`. ```python from collections import Counter class WSBEvent(BaseEvent): ... class WSBSentimentEvent(BaseEvent): ... class WSJEvent(BaseEvent): ... class WSJSentimentEvent(BaseEvent): ... class TrendingPredictionEvent(BaseEvent): ... class PredictionResultEvent(BaseEvent): ... class TradeEvent(BaseEvent): ... ``` Then, we write a `ComplexTradeMonitor` class as a new Workflow. ```python class ComplexTradeMonitor(Workflow): def __init__(self, *args, **kwargs): self.trader = StockTrader() super().__init__(*args, **kwargs) @step async def start(self, ctx: Context, ev: StartEvent) \ -> WSBEvent | WSJEvent | TrendingPredictionEvent: self.stock = ev.payload ctx.send_event(WSBEvent(payload=ev.payload)) ctx.send_event(WSJEvent(payload=ev.payload)) ctx.send_event(TrendingPredictionEvent(payload=ev.payload)) @step async def wsb_sentiment(self, ev: WSBEvent) -> WSBSentimentEvent: await asyncio.sleep(random.randint(1, 3)) if random.random() < 0.3: return WSBSentimentEvent(payload='Bearish') else: return WSBSentimentEvent(payload='Bullish') @step async def wsj_sentiment(self, ev: WSJEvent) -> WSJSentimentEvent: await asyncio.sleep(random.randint(1, 3)) if random.random() < 0.3: return WSJSentimentEvent(payload='Bearish') else: return WSJSentimentEvent(payload='Bullish') @step async def trending_predict(self, ev: TrendingPredictionEvent) -> PredictionResultEvent: await asyncio.sleep(random.randint(1, 3)) if random.random() < 0.3: return PredictionResultEvent(payload='Bearish') else: return PredictionResultEvent(payload='Bullish') @step async def trading_decision(self, ctx: Context, ev: WSBSentimentEvent | WSJSentimentEvent | PredictionResultEvent)\ -> TradeEvent: results = ctx.collect_events(ev, [WSBSentimentEvent, WSJSentimentEvent, PredictionResultEvent]) if results is not None: voting = dict(Counter([ev.payload for ev in results])) print(voting) sentiment = max(voting, key=voting.get) return TradeEvent(payload=sentiment) @step async def trade(self, ev: TradeEvent) -> StopEvent: sentiment = ev.payload match sentiment: case 'Bullish': await self.trader.buy(stock=self.stock) case 'Bearish': await self.trader.sell(stock=self.stock) case _: print("Do nothing") return StopEvent(result='We are done for the day.') ``` In the `start` method, we use `ctx.send_event` to simultaneously throw out `WSBEvent`, `WSJEvent`, and `TrendingPredictionEvent`. Since Workflow determines which messages were thrown out based on the typing annotation of the `step` method, we still need to mark the returned message types. Next, we implement the `wsb_sentiment`, `wsj_sentiment`, and `trending_predict` methods to obtain trading signals and return the corresponding events. We still use the `random` method to simulate market sentiment analysis. As content from different sources requires different parsing times, we hope to wait until all messages return before making a trading decision. At this point, we can use the `ctx.collect_events` method in the `trading_decision` method. Each time a new sentiment event returns, the `trading_events` method executes once. But the `ctx.collect_events` method takes all the events we need to wait for as parameters, and its return value remains empty until all sentiment events return. At that point, the return value is a list of three sentiment events. We can use the `Counter` method to count how many times Bullish and Bearish appear, then take the most voted mark to make a trading decision. Finally, let's use the `draw_all_possible_flows` tool to see how cool our newly designed workflow is: ```python draw_all_possible_flows(ComplexTradeMonitor, filename='complex_trade_monitor.html') ``` ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-5.png) Workflow executes three asynchronous tasks in parallel and gets the final result. Image by Author Next, let's execute this workflow and see. ```python trade_monitor = ComplexTradeMonitor(timeout=20, verbose=True) result = await trade_monitor.run(payload='[Stock Code]') print(result) ``` ![Detailed process of executing code in parallel with Workflow.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-6.png) Detailed process of executing code in parallel with Workflow. Image by Author We can observe that the three methods to obtain sentiment from different sources are triggered simultaneously but return at different times. The first two returned events can trigger the `trading_decision` method, but cannot continue to trigger the `TradeEvent`. Only after all three events return and the final trading decision is calculated, is the `TradeEvent` triggered. As you can see, with the power of Workflow, we can indeed make our code architecture both clear and efficient. But don't be too optimistic, because after some time in practice, I think there are still some shortcomings. --- ## Time to Talk about Shortcomings If you review our previous code, you'll notice that all our code logic is written in the same Workflow, which is fine for simple applications but a disaster for complex real-world applications. Ideally, we should split different logic into Workflows to maintain the purity of the "single responsibility" principle. The official solution to this requirement is [nested workflows](https://docs.llamaindex.ai/en/stable/understanding/workflows/nested/?ref=dataleadsfuture.com): ### Nested workflows Suppose we want to split the trading order logic from the `TradeMonitor` into an independent Workflow. How should we call it when we need to place an order? The official solution is a nested workflows, i.e., passing another workflow B as a parameter in the `step` method of workflow A. Then, after workflow A is instantiated, add the instance of workflow B. As shown in the following code: ```python class TradeStation(Workflow): def __init__(self, *args, **kwargs): self.trader = StockTrader() super().__init__(*args, **kwargs) @step async def trade(self, ev: StartEvent) -> StopEvent: print("We are now in a new workflow named TradeStation") sentiment = ev.sentiment match sentiment: case 'Bullish': await self.trader.buy(stock=ev.stock) case 'Bearish': await self.trader.sell(stock=ev.stock) return StopEvent(result="Done!") class ComplexTradeMonitorV2(ComplexTradeMonitor): @step async def trade(self, ev: TradeEvent, trade_station: TradeStation) -> StopEvent: sentiment = ev.payload await trade_station.run(sentiment=sentiment, stock=self.stock) return StopEvent(result='We are done for the day.') ``` ```python trade_monitor_v2 = ComplexTradeMonitorV2(timeout=20, verbose=False) trade_monitor_v2.add_workflows( trade_station=TradeStation(timeout=10, verbose=True) ) result = await trade_monitor_v2.run(payload='[Stock Code]') print(result) ``` Wait a minute, if you have Java development experience, will you be surprised to see this code: isn't this dependency injection? ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-nested-workflow.drawio.png) Nested Workflow works like “dependency injection”. Image by Author It's indeed similar to dependency injection, but the difference is that we still need to explicitly add the specific workflow instance after the instance is initialized, so there is still coupling, which is the first problem. Another problem I found during coding is that for nested workflows, I can only call them through the `run` method, not by calling the corresponding `step` method in the nested workflow from an external workflow. Therefore, this is not a good solution for communication between workflows. ### Communicate between Workflows So, is there any way to truly achieve communication between workflows? I searched the API documentation and couldn't find an official solution, and I noticed that this [issue](https://github.com/run-llama/llama%5Findex/issues/15466?ref=dataleadsfuture.com) also went unanswered. So I decided to try it myself to see if I could solve it. After reviewing the source code again, I think the `ctx.send_event` method has some potential, so the first thing I thought of was whether sharing the same Context between two workflows could solve it? I noticed that instantiating `Context` requires passing in a `workflow` instance and setting a workflow's own Context can be done by passing it in during the `run` method. So the code is as follows, keeping the two workflows unchanged, only that the `step` method in the `TradeStation` no longer accepts a `StartEvent` but a specific `TradeEventV2`. ```python class TradeEventV2(Event): sentiment: str stock: str class TradeStation(Workflow): def __init__(self, *args, **kwargs): self.trader = StockTrader() super().__init__(*args, **kwargs) @step async def trade(self, ev: TradeEventV2) -> StopEvent: print("We are now in a new workflow named TradeStation") sentiment = ev.sentiment match sentiment: case 'Bullish': await self.trader.buy(stock=ev.stock) case 'Bearish': await self.trader.sell(stock=ev.stock) return StopEvent(result="Done!") ``` Then I use the `TradeStation` to create a Context instance and pass it into the `TradeMonitor` instance during the run method execution, and sure enough, it throws an error: ```python trade_monitor_v3 = ComplexTradeMonitorV3(timeout=20, verbose=False, disable_validation=False) trade_station = TradeStation(timeout=10, verbose=True) result = await trade_monitor_v3.run(ctx=Context(workflow=trade_station), payload='[Stock Code]') print(result) ``` ![I get an error when I use Context to communicate between two Workflows.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-7.png) I get an error when I use Context to communicate between two Workflows. Image by Author It seems there is a problem with the method signature validation, let's try turning off the validation: ```python trade_monitor_v3 = ComplexTradeMonitorV3(timeout=20, verbose=False, disable_validation=True) trade_station = TradeStation(timeout=10, verbose=True) result = await trade_monitor_v3.run(ctx=Context(workflow=trade_station), payload='[Stock Code]') print(result) ``` Still no luck, it seems this way won't work. ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/image-8.png) The trade method in TradeStation is not triggered. Image by Author ### Unbound syntax Then, I noticed that the documentation mentioned a kind of [Unbound syntax](https://docs.llamaindex.ai/en/stable/understanding/workflows/unbound%5Ffunctions/?ref=dataleadsfuture.com), which seems to be able to decouple each step's logic from the Workflow. The example code is as follows: ```python class TestWorkflow(Workflow): ... @step(workflow=TestWorkflow) def some_step(ev: StartEvent) -> StopEvent: return StopEvent() ``` Although we can still only run within one Workflow, it made me feel the feasibility of communication between modules. Due to the length of the article, I won't use code to explain here, let me show you a diagram of how to use Unbound syntax for module communication: ![A diagram to describe how Unbound syntax decouples code logic into multiple modules. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/12/llamaindex_workflow-Unbound-Syntax.drawio.png) A diagram to describe how Unbound syntax decouples code logic into multiple modules. Image by Author As shown in the diagram: First, we can define an `Application` class as a Workflow pipeline, and simultaneously define the required events. Then, each project team can write their own business logic code and use different `step` methods to listen and send messages externally. Finally, we can call the `run` method of the `Application` in the fastapi API to mobilize various modules to complete the task. In this way, business logic can be split into different modules for development, and then different `step` methods can be called using events. This indeed achieves the purpose of logic decoupling. However since this method only registers each step to the Workflow in the `step` decorator through the `add_step` method, it still does not achieve real communication between Workflows. --- ## Summary The new feature of LlamaIndex's Workflow, makes parallel execution of RAG, LLM generation, and I/O calls a very simple task, and the event-driven architecture also allows the program to decouple from complex logic control. In today's article, I demonstrated several features of Workflow through a TradeMonitor project. In project practice, we also found that Workflow still has shortcomings in communication between modules, and we discussed different solutions including nested workflows and unbound syntax. Finally, as agent frameworks like Langchain and AutoGen start to propose their own event-driven architectures, I believe Workflow is on the right path and will see long-term development. Let's keep an eye on it. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- next step Mr. Qian's Recommendation: Designing an event-driven LlamaIndex Workflow is one of the biggest steps you can take. From here, you should think about how to use LlamaIndex to build a production-ready, high-accuracy RAG system. We recommend checking out the [****RAG for Generative AI Applications specialization**](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) course brought to you by industry experts at IBM. You'll get hands-on experience with all the key tools and skills you need to build a RAG Workflow, so you can hit the ground running from day one. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [Start Learning for Free Now ](https://imp.i384100.net/GbQnNk?ref=dataleadsfuture.com) --- ## Further Reading Next, we will discuss how to use the LlamaIndex Workflow to achieve agent handoff capabilities similar to OpenAI Swarm: [Using LLamaIndex Workflow to Implement an Agent Handoff Feature Like OpenAI SwarmExample: a customer service chatbot project![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color-icon-1-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-1-1.webp)](https://www.dataleadsfuture.com/using-llamaindex-workflow-to-implement-an-agent-handoff-feature-like-openai-swarm/) Still feeling confused about how to use Workflow? No worries, LlamaIndex introduced AgentWorkflow in early '25\. It's the super-evolved version of Workflow—simple, easy to learn, and almost perfect: [Diving into LlamaIndex AgentWorkflow: A Nearly Perfect Multi-Agent Orchestration SolutionAnd fix the issue where the agent can’t continue with past requests![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-21.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-4-1-1.webp)](https://www.dataleadsfuture.com/diving-into-llamaindex-agentworkflow-a-nearly-perfect-multi-agent-orchestration-solution/) ### How to Become a Data Scientist If You Lacking the Necessary Skills URL: https://www.dataleadsfuture.com/how-to-become-a-data-scientist-if-you-lacking-the-necessary-skills/ Last updated: 2025-03-26T02:17:56.000Z Recently, during a lecture for campus recruitment, a student asked me: He only knows math and basic sciences but doesn't know how to code. Can he still become a data scientist? First off, yes, absolutely. How to do it? First, understand if becoming a data scientist is your goal or a part of your journey. From what I see, you're asking how to achieve this goal. That involves figuring out what skills you need and which paths to take. --- ## Start Working Hard Now No one understands all knowledge areas from the start. Generally, school education lays a foundation for your career but that doesn't mean your learning stops at graduation. Academia often lags behind market demands. We can't all start with the skills needed for a job. Even if we have them, they're just formulas and theories from books, far from what companies actually need. Take me, for example. I've been in the big data field for nearly 15 years, working as a senior data scientist at a giant internet company, and now I'm the chief data architect at a large brokerage firm's big data department. But you know what? When I was about to graduate, I was just an undergrad in business administration from an ordinary university. Unlike you, I started with some programming experience. I had taken courses in linear algebra and calculus, but at the time, I didn't see the point in learning these theories. My first challenge was finding a decent job, not thinking about a long-term goal like data science. So, I took stock of my strengths and available tools. I loved programming; I had plenty of time, unlike now, and I could access the school's library and computers for free, giving me a low-cost learning opportunity. While my dorm mates were lost in World of Warcraft, I spent two hours every day immersed in programming. I mastered basic design patterns and algorithms and even got a Java certification. Also, I'm an introvert, or a nerd as you might say, not keen on talking much. So, I forced myself to join student council activities and interact with opinion leaders or join some speeches to develop my ability to express myself. Looking back, if I could do college over, I'd take more internships to better understand what companies were doing to prepare faster. Is it tiring? Yes, but we have no choice. Once we set a "big" goal, we must keep going. Forced persistence won't work; you must also love what you do. --- ## Keep the Love, Stay Focused Common wisdom says when you graduate, you're a blank slate with endless possibilities and many potential paths. But why choose to be a data scientist? Apart from the recent boom in big AI models, maybe because you love the field. Love is the original driving force that keeps a person growing and persisting. Why did I want to be in programming? In college, I deeply considered my future. Most from my business admin program went into real estate sales or maybe design institutes if they were more skilled. Could I handle those jobs? The thought of dealing daily with clients or attending lengthy lunches made me feel unfit. What could I do? I remember programming in BASIC in elementary school, writing scripts in Flash during high school, and doing well in college programming courses. Though these were mere toys compared to enterprise programming, I genuinely enjoyed sitting at a computer with reference docs all afternoon and not getting bored. I think that was proof of my love for the field. So why consider other possibilities? I decided to stick with programming and have been in the field ever since, despite temptations and confusion along the way. The 10,000-hour rule is real. If you keep loving and pursuing this field, you'll become an expert. Okay, I've said a lot that seems unrelated to data science. Can we get back on topic? Data science wasn't a job when I started working; the closest was "data mining engineer." As companies grew and technology advanced, data science emerged. --- ## Embrace Change You never cross the same river twice. Likewise, business logic won't wait for you to get ready. To stay competitive, you must continually adapt your skills to meet talent demands. I'm not an HR expert, but I'll share how I grew into a data scientist. After graduating, thanks to my schooling, I joined a giant company's data mining team as a junior developer. Initially, I did data development. As the team grew, we needed Java engineers for web apps, and I was chosen despite my reluctance since I was just getting comfortable with SQL. But this change was a blessing. Working as a Java engineer let me apply the design patterns and object-oriented programming I learned in college. Being closer to business in app development helped me understand end-user pain points, aiding my career growth. Over the years, I switched roles as needed: from enhancing user experience as a front-end developer to building data platforms back in data development, and managing recommendation models as part of MLOps, eventually becoming a data science expert. Some ask, didn't you say to stay focused? Why change roles? I don't see a contradiction. When you extend your career over decades, it's hard to stick to one niche. By exploring different roles, I understood what companies needed from an engineer. I stayed in big data, and my diverse background helped me communicate and collaborate across roles. But no matter what, you're working to earn your salary. How do you excel in these various roles? My answer is continuous learning. --- ## Continuous learning Initially, I had basic Java skills, but my first job was writing SQL. Imagine my excitement and fear when I got the offer. I couldn't afford to lose this job, so I spent a week learning Oracle database development, sleeping only a few hours a night. I don't advocate pressuring yourself at work, but for a career you love, it's worth the effort. During my career, I often took time to follow the latest tech trends and fill my skill gaps. Opportunities come to those who are prepared. I didn't wait to need a job to learn its required skills. Because I was interested, I prepared in advance, winning the chance to work those jobs. This leads to today's last topic: How can I learn more effectively? My answer: actively practice. --- ## Active Practice There's a Chinese saying: "A good memory is not as good as a bad pen." Rather than relying on memory for everything you read, it's better to write summaries and solve problems. If you're into Kong Fu, you know that just practicing a manual from start to finish is "by the book" learning. You need practical experience to handle various situations. So, get involved in work to find your or your users' pain points, then learn everything to solve them. If your job lacks practical opportunities, I recommend participating in projects outside of work. Joining open-source projects is a great way to let more people see your work and help you improve. Or you could join Kaggle competitions, offering low-cost trial-and-error opportunities. Or like me, start a blog like Data Leads Future to document your projects, summarize your progress, and help others. --- ## Summary Let's return to the beginning. How do you land your dream job without specific skills? My answer: Start working hard now, keep loving and focusing, embrace change, continue learning, and actively practice. Becoming a data scientist might seem amazing now, but over your career, it's just one milestone. Once you build enough strength, I believe you'll achieve greater success and make more contributions. ### How to Connect LlamaIndex with Private LLM API Deployments URL: https://www.dataleadsfuture.com/how-to-connect-llamaindex-with-private-llm-api-deployments/ Last updated: 2025-05-23T03:16:54.000Z ## Introduction Starting with LlamaIndex is a great choice when building an RAG pipeline. Usually, you need an OpenAI API key to follow the many tutorials available. However, you might face these situations: - Your company can only use privately deployed models due to compliance. - You're using a model fine-tuned by your data science team. - Legal restrictions prevent your company's private data from leaving the country. - Other reasons that require using privately deployed models. When building enterprise AI apps, you can't use OpenAI or other cloud providers' LLM services. This leads to a frustrating first step: How do I connect my LlamaIndex code to my company's private API service? To save your time, if you just need the solution, install this extension: ```Bash pip install -U llama-index-llms-openai-like ``` This will solve your problem. If you want to understand why, let's continue. --- ## Why Install an Extra SDK? LlamaIndex uses OpenAI as the default LLM. Here's a sample code from their website using DeepLake as the vector store: ```python from llama_index.core import VectorStoreIndex, SimpleDirectoryReader documents = SimpleDirectoryReader("./data/").load_data() vector_store_index = VectorStoreIndex.from_documents(documents) vector_query_engine = vector_store_index.as_query_engine(similarity_top_k=k, temperature=temp, num_output=mt) def index_query(input_query: str) -> tuple: response = vector_query_engine.query(input_query) print("LLM successfully generated the desired content") index_query(user_input) ``` I've removed other code parts - this is just for the demo. You can find the complete code on [their website](https://docs.llamaindex.ai/en/stable/examples/vector%5Fstores/DeepLakeIndexDemo/?ref=dataleadsfuture.com). If you haven't subscribed to OpenAI, you'll get this error: ```bash NotFoundError: Error code: 404 - {'error': {'message': 'The model `gpt-3.5-turbo` does not exist or you do not have access to it.', 'type': 'invalid_request_error', 'param': None, 'code': 'model_not_found'}, 'request_id': '26a1b8f0-10b3-9e03-9064-0f5320995f3d'} ``` ![If you haven't subscribed to OpenAI, you'll get this error.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/11/image.png) If you haven't subscribed to OpenAI, you'll get this error. Image by Author This is the first error you'll face when deploying to production. Let's see how to fix this when using a private `Qwen-max` service. --- ## Try LlamaIndex's OpenAI Class Our MLops team prefers tools like `vllm` or `sglang` to deploy Qwen services. For compatibility, Qwen models support OpenAI's API format. This sounds good. Let's check the docs to see if we can use LlamaIndex's `OpenAI` class. Note: To specify which LLM class to use, set the `Settings.llm` property. ```python from llama_index.llms.openai import OpenAI Settings.llm = OpenAI( model="qwen-max", is_chat_model=True ) ``` You must set `OPENAI_API_KEY` and `OPENAI_API_BASE` environment variables to your company's values. Let's try running the code: ```bash ValueError: Unknown model 'qwen-max'. Please provide a valid OpenAI model name in: o1-preview, o1-preview-2024-09-12, o1-mini, o1-mini-2024-09-12, gpt-4, gpt-4-32k, gpt-4-1106-preview, gpt-4-0125-preview, gpt-4-turbo-preview, gpt-4-vision-preview, gpt-4-1106-vision-preview, gpt-4-turbo-2024-04-09, gpt-4-turbo, gpt-4o, gpt-4o-2024-05-13, gpt-4o-2024-08-06, chatgpt-4o-latest, gpt-4o-mini, gpt-4o-mini-2024-07-18, gpt-4-0613, gpt-4-32k-0613, gpt-4-0314, gpt-4-32k-0314, gpt-3.5-turbo, gpt-3.5-turbo-16k, gpt-3.5-turbo-0125, gpt-3.5-turbo-1106, gpt-3.5-turbo-0613, gpt-3.5-turbo-16k-0613, gpt-3.5-turbo-0301, text-davinci-003, text-davinci-002, gpt-3.5-turbo-instruct, text-ada-001, text-babbage-001, text-curie-001, ada, babbage, curie, davinci, gpt-35-turbo-16k, gpt-35-turbo, gpt-35-turbo-0125, gpt-35-turbo-1106, gpt-35-turbo-0613, gpt-35-turbo-16k-0613 ``` ![I pointed OPENAI_API_BASE to our company endpoint but still got GPT-related errors.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/11/image-1.png) I pointed `OPENAI_API_BASE` to our company endpoint but still got GPT-related errors. Image by Author Strange - I pointed `OPENAI_API_BASE` to our company endpoint but still got GPT-related errors. Unlike LangChain, LlamaIndex's `OpenAI` class checks the `model_name` in `metadata` to handle different model features. It forces you to use GPT family models. So this class won't work with other models. --- ## Try OpenAILike Instead As mentioned earlier, we can use the `openai-like` extension to connect to our API service. Let's read [the API docs](https://docs.llamaindex.ai/en/stable/api%5Freference/llms/openai%5Flike/?ref=dataleadsfuture.com) (which are quite hidden): The docs say: > "OpenAILike is a thin wrapper around the OpenAI model that makes it compatible with 3rd party tools that provide an openai-compatible API. > Currently, llama\_index prevents using custom models with their OpenAI class because they need to be able to infer some metadata from the model name." This explains why we can't use custom models with the `OpenAI` class, and why `OpenAILike` solves the problem. Let's update our code and try again: ```python from llama_index.llms.openai_like import OpenAILike Settings.llm = OpenAILike( model="qwen-max", is_chat_model=True ) ``` **You must set the `is_chat_model` parameter to `True`, otherwise, you'll get a 404 error. Or there might be other random issues.** If you need to use the `function_calling` API or develop an agent, you need to set `is_function_calling_model=True`. ![LLM API services has been successfully connected.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/11/image-2.png) LLM API services has been successfully connected. Image by Author Bingo - no errors and the LLM is connected. I've written another article that thoroughly explains the integration solutions for privately deployed LLMs, along with support for some new features. You can click here to learn more: [Build AutoGen Agents with Qwen3: Structured Output & Thinking ModeSave yourself 40 hours of trial and error![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-32.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Autogen_Qwen3_cover_2.webp)](https://www.dataleadsfuture.com/build-autogen-agents-with-qwen3-structured-output-thinking-mode/) --- ## Conclusion When trying to use LlamaIndex in enterprise RAG pipelines, I struggled to connect to private LLM services. Despite lots of Googling, no tutorial explained how to solve this. I had to dig through LlamaIndex's API docs to find the answer. That's why I wrote this short article - to help you solve this quickly. My team is just starting to build LLM apps in finance. I hope to discuss various challenges with you. Feel free to leave comments - I'll reply soon. --- Next, we will embark on a journey to explore the LlamaIndex Workflow: [Deep Dive into LlamaIndex Workflow: Event-driven LLM architectureWhat I think about the progress and shortcomings after practice![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color-icon-1-2.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/LLamaIndex_Workflow-1.webp)](https://www.dataleadsfuture.com/deep-diving-into-llamaindex-workflow-event-driven-llm-architecture/) [Using LLamaIndex Workflow to Implement an Agent Handoff Feature Like OpenAI SwarmExample: a customer service chatbot project![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color-icon-1-3.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/cover-1.webp)](https://www.dataleadsfuture.com/using-llamaindex-workflow-to-implement-an-agent-handoff-feature-like-openai-swarm/) ### Implementing the train_step Method in Keras 3: From Errors to Solutions URL: https://www.dataleadsfuture.com/implementing-the-train_step-method-in-keras-3-from-errors-to-solutions/ Last updated: 2025-03-26T02:17:19.000Z Keras has been updated to version 3.0, but many code examples on the official website have not been maintained in time. So, when you use Pytorch as the backend framework, these code examples will likely fail. For instance, take the example related to the [VAE](https://keras.io/examples/generative/vae/?ref=dataleadsfuture.com) model, and another concerning the [GAN](https://keras.io/examples/generative/dcgan%5Foverriding%5Ftrain%5Fstep/?ref=dataleadsfuture.com) model. Both examples share a common feature: they rewrite the `train_step` method of `keras.Model` to implement a custom model training and gradient update process. However, following these examples will likely result in errors. After several failed attempts, I finally figured out how to write the `train_step` method correctly. Today, I will share my solution with you, hoping to help you solve similar problems. If you are unfamiliar with the new changes in Keras 3, you can read my deep dive article here: [Keras 3.0 Tutorial: End-to-End Deep Learning Project GuideImplement an encoder-decoder recurrent network from scratch![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2024/02/color-icon-1.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/05/keras3.0_encoder_decoder-1.webp)](https://www.dataleadsfuture.com/keras-3-0-tutorial-end-to-end-deep-learning-project-guide/) --- ## Introducing Some Background Knowledge Before we start, I need to provide you with some background knowledge about the Keras training loop, as shown in the diagram below. ![The training loop in Keras.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/09/keras_train_step.drawio.png) The training loop in Keras. Image by Author In the `fit` method of `keras.Model`, you will set two parameters: `epochs` and `batch_size`. `Epochs` set how many times we train the model with the entire data. For one training loop, we use `batch_size` to divide the data into multiple batches. Each batch is a step, and the model only starts gradient updates after a step is trained. For each step, Keras divides it into `train_step` and `test_step` methods. The `train_step` method is executed only during training, while the `test_step` method is performed during validation. The `train_step` method, includes four processes: forward propagation, loss calculation, gradient update, and updating metrics. So when we need to customize the model training process, we can implement it in the `train_step` method. --- ## Getting to the Point ### Starting with a base model I don't like to start by throwing a whole project's code at you, as it can be overwhelming and unclear where the focus should be. So let's keep it simple, starting with a basic model based on the functional API. ```python (X_train, y_train), (X_test, y_test) = keras.datasets.mnist.load_data() X_train = X_train.astype('float32') / 255. X_test = X_test.astype('float32') / 255. def get_base_model() -> keras.models.Model: inputs = layers.Input(shape=(28, 28, 1)) x = layers.Flatten()(inputs) x = layers.Dense(128, activation='relu')(x) x = layers.Dense(64, activation='relu')(x) x = layers.Dropout(0.25)(x) outputs = layers.Dense(10, activation='softmax')(x) return keras.models.Model(inputs=inputs, outputs=outputs) base_model = get_base_model() base_model.compile(loss='sparse_categorical_crossentropy', optimizer='adam', metrics=['accuracy']) base_model.fit(X_train, y_train, epochs=5, batch_size=512, validation_data=(X_test, y_test)) ``` This model is a very simple task of classifying the MNIST dataset. So the training process is also straightforward: set the `optimizer`, `loss`, and `metrics` in the `compile` method, and you can start training. ### Let's see how the functional API works First, let's add some custom processes, assuming we do not use the `loss` set in the compile method, but instead want to customize the loss calculation and gradient update methods. We need to use the subclassing API to implement a Model subclass and then rewrite the `train_step` method. ```python class CustomModel(keras.Model): def __init__(self, *args, **kwargs): super().__init__(*args, **kwargs) self.loss_fn = keras.losses.SparseCategoricalCrossentropy() self.total_loss_tracker = keras.metrics.Mean(name='total_loss') @property def metrics(self): return [ self.total_loss_tracker ] def train_step(self, data): X, y = data self.zero_grad() y_pred = self(X, training=True) total_loss = self.loss_fn(y, y_pred) total_loss.backward() trainable_variables = self.trainable_variables gradients = [v.value.grad for v in trainable_variables] with torch.no_grad(): self.optimizer.apply(gradients, trainable_variables) self.total_loss_tracker.update_state(total_loss) return { 'custom_loss' : self.total_loss_tracker.result() } ``` We focus on demonstrating the steps of `train_step`, so inside the sub-Model, I still use `SparseCategoricalCrossentropy` to calculate the loss. At the same time, we need to replace the Model class used by the functional API. ```python def get_custom_model(): inputs = layers.Input(shape=(28, 28, 1)) x = layers.Flatten()(inputs) x = layers.Dense(128, activation='relu')(x) x = layers.Dense(64, activation='relu')(x) x = layers.Dropout(0.25)(x) outputs = layers.Dense(10, activation='softmax')(x) return CustomModel(inputs=inputs, outputs=outputs) ``` Since we are using Pytorch as the backend framework, I will also show you what the gradient update process based on Pytorch looks like. As shown in the code, we need to pay attention to the following lines of code: 1. When using Pytorch as the backend, the loss calculation process no longer needs to be placed in the `tf.GradientTape()` context. 2. Before each batch starts, we need to use `self.zero_grad()` to reset the gradients. 3. After the loss calculation is completed, `loss.backward()` will perform the backward propagation process. 4. We need to perform gradient updates in the `torch.no_grad()` context to avoid repeated gradient calculations. If we compare our code with the official website's example, we've done quite well. Next, let's remove the `loss` and `metrics` parameters from the `compile` method (since we have already manually calculated them in the sub-Model), and then try training the model. ```python custom_model = get_custom_model() custom_model.compile(optimizer='adam') custom_model.fit(X_train, y_train, epochs=5, batch_size=512, validation_data=(X_test, y_test)) ``` Boom, the training went wrong, facing our first error today. ```bash ValueError: No loss to compute. Provide a `loss` argument in `compile()`. ``` The error message is hard to understand: We have already calculated the loss in the `train_step` method, so why do we still need to provide the `loss` parameter in the `compile` method? Where is this `loss` method used? Recalling the model training process diagram mentioned earlier, I noted that when the `fit` method includes `validation_data` parameter, we need to rewrite the `test_step` method to indicate the calculation process of the validation data. So we need to rewrite another `test_step` method in the sub-Model, which is quite simple, just writing out the forward propagation calculation process. ```python def test_step(self, data): X, y = data y_pred = self.model(X, training=False) for metric in self.metrics: metric.update_state(y, y_pred) return {m.name: m.result() for m in self.metrics} ``` Of course, for models like GANs, where the forward propagation itself is quite complex, it's not easy to implement in the `test_step`. We can also directly remove the `validation_data` or `validation_split` parameters from the `fit` method. ### How will it work with the subclassing API We talked about the scenario of the functional API earlier. Now let's talk about what problems might arise when rewriting the `train_step` under the subclassing API, and how to solve them. Typically, when our training process is too complex to be implemented with the functional API, we recommend using the subclassing API. We can take the AutoEncoder architecture as an example. ![A typical AutoEncoder architecture diagram.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/09/autoencoder.drawio.png) A typical AutoEncoder architecture diagram. Image by Author Suppose we rewrite the code for MNIST classification into an Encoder. ```python class Encoder(keras.Model): def __init__(self, **kwargs): super().__init__(**kwargs) self.flatten = layers.Flatten() self.fc1 = layers.Dense(128, activation='relu') self.fc2 = layers.Dense(64, activation='relu') self.dropout = layers.Dropout(0.25) self.out = layers.Dense(10, activation='softmax') def call(self, data): x = self.flatten(data) x = self.fc1(x) x = self.fc2(x) x = self.dropout(x) out = self.out(x) return out ``` Then we put it into an AutoEncoder, ready to customize the training process. To simplify the code, we skip the implementation and training process of the Decoder. The implementation of the `train_step` method is consistent with the previous CustomModel. Next, let's see what happens if we train the AutoEncoder. ```bash RuntimeError: Unable to automatically build the model. Please build it yourself before calling fit/evaluate/predict. A model is 'built' when its variables have been created and its `self.built` attribute is True. Usually, calling the model on a batch of data is the right way to build it. Exception encountered: 'Exception encountered when calling AutoEncoder.call(). Model AutoEncoder does not have a `call()` method implemented. Arguments received by AutoEncoder.call(): • args=('torch.Tensor(shape=torch.Size([512, 28, 28]), dtype=float32)',) • kwargs=' ``` Another error! The error text is a bit long, but the key point is in the middle: ```bash Model AutoEncoder does not have a `call()` method implemented. ``` Wait, what? If you look at the official website's several examples about the `train_step`: [Example 1](https://keras.io/guides/custom%5Ftrain%5Fstep%5Fin%5Ftorch/?ref=dataleadsfuture.com), [Example 2](https://keras.io/examples/generative/vae/?ref=dataleadsfuture.com), [Example 3](https://keras.io/examples/generative/dcgan%5Foverriding%5Ftrain%5Fstep/?ref=dataleadsfuture.com), [Example 4](https://keras.io/examples/keras%5Frecipes/trainer%5Fpattern/?ref=dataleadsfuture.com). They didn't mention that when rewriting the `train_step` loop, you need to implement the `call` method. And for models like VAE or GAN, where the training process is quite complex, there isn't a need to implement the `call` method. So why are we still prompted to implement the `call` method here? After checking the Keras source code and comparing the differences between TensorFlow and Pytorch, I have a hypothesis: This may be related to Keras's computation graph. We all know that when using Keras's Sequential and functional API to build models, a static computation graph is generated, but not in the case of using the subclassing API. Therefore, for a subclassing API Model, it needs to use `call` to determine the computational architecture between the layers in the model. So when we rewrite the `train_step` using the subclassing method, it prompts the *"does not have a call() method implemented."* error. Based on this hypothesis, let's add a `call` method to the AutoEncoder and try again. ```python def call(self, data): x = self.encoder(data) return x ``` Bingo, the model training starts normally. ```bash Epoch 1/5 118/118 ━━━━━━━━━━━━━━━━━━━━ 1s 10ms/step - total_loss: 1.0065 Epoch 2/5 118/118 ━━━━━━━━━━━━━━━━━━━━ 1s 8ms/step - total_loss: 0.2299 Epoch 3/5 118/118 ━━━━━━━━━━━━━━━━━━━━ 1s 8ms/step - total_loss: 0.1600 Epoch 4/5 118/118 ━━━━━━━━━━━━━━━━━━━━ 1s 8ms/step - total_loss: 0.1270 Epoch 5/5 118/118 ━━━━━━━━━━━━━━━━━━━━ 1s 8ms/step - total_loss: 0.1013 ``` Similarly, when you learn from the official website's examples of VAE and GAN and encounter code errors, you can also solve them by adding a `call` method. --- ## Conclusion Due to the scarcity of related materials, when we learn about Keras 3, we mainly rely on the official website's examples. However, these examples seem not to have kept up with the framework updates, causing some code to report errors when executed with Pytorch as the backend. This article focuses on the feature of implementing a custom training process by rewriting the `train_step` method, explaining the possible errors and solutions. Of course, if a simpler solution is available to train the model, it is not recommended to implement the training process yourself, as debugging deep learning is quite troublesome and often encounters various problems. If you want to learn more about Keras 3, feel free to leave me a message, and I will do my best to answer. ### The Math Behind Keras 3 Optimizers: Deep Understanding and Application URL: https://www.dataleadsfuture.com/the-math-behind-keras-3-optimizers-deep-understanding-and-application/ Last updated: 2025-03-26T02:16:54.000Z ## Introduction Optimizers are an essential part of everyone working in machine learning. We all know optimizers determine how the model will converge the loss function during gradient descent. Thus, using the right optimizer can boost the performance and the efficiency of model training. Besides classic papers, many books explain the principles behind optimizers in simple terms. However, I recently found that the performance of Keras 3 optimizers doesn't quite match the mathematical algorithms described in these books, which made me a bit anxious. I worried about misunderstanding something or about updates in the latest version of Keras affecting the optimizers. So, I reviewed the source code of several common optimizers in Keras 3 and revisited their use cases. Now, I want to share this knowledge to save you time and help you master Keras 3 optimizers more quickly. If you're not very familiar with the latest changes in Keras 3, here's a quick rundown: Keras 3 integrates TensorFlow, PyTorch, and JAX, allowing us to use cutting-edge deep learning frameworks easily through Keras APIs. [Keras 3.0 Tutorial: End-to-End Deep Learning Project GuideImplement an encoder-decoder recurrent network from scratch![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2024/02/color-icon-1.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/05/keras3.0_encoder_decoder-1.webp)](https://www.dataleadsfuture.com/keras-3-0-tutorial-end-to-end-deep-learning-project-guide/) --- ## **Preparation** ### Some utility methods I plan to show you through some charts how these optimizers affect the convergence of loss functions. Before starting, I need to prepare some utility methods for creating these charts. First, I'll use sklearn to generate a virtual dataset for a classification task: ```python from sklearn.datasets import make_moons X, y = make_moons(n_samples=300, noise=0.1, random_state=42) ``` Since the dataset is quite simple, I plan to build a basic multilayer perceptron model with three hidden layers and a dual-node output layer: ```python from keras import layers, utils, ops import keras def build_model(input_shape: tuple): inputs = layers.Input(shape=input_shape) x = layers.Dense(12, activation='relu')(inputs) x = layers.Dense(13, activation='relu')(x) x = layers.Dense(13, activation='relu')(x) outputs = layers.Dense(2, activation='softmax')(x) return keras.Model(inputs=inputs, outputs=outputs) ``` Finally, our utility method will evaluate how different optimizers affect model loss convergence. It also takes a parameter for epochs to highlight details of some optimizers. This method will also use the matplotlib library to plot the model loss convergence curves, making the assessment more visual: ```python def fit_show_model(optimizer: keras.optimizers.Optimizer, epochs: int = 800 ): my_model = build_model(input_shape=(2,)) my_model.compile(optimizer=optimizer, loss='sparse_categorical_crossentropy', metrics=['accuracy']) history = my_model.fit(X, y, batch_size=32, epochs=epochs, verbose=0) loss=history.history['loss'] _, ax = plt.subplots(figsize=(5, 3)) ax.plot(range(len(loss)), loss, 'b') ax.set( xlabel="Epoch", ylabel="Loss" ) plt.show() ``` ### Variable abbreviations Since this article is full of mathematical expressions, I plan to abbreviate some common variables to make these expressions clearer. For example: - `lr` stands for `learning_rate`. - `g` represents the gradient at the current node. - `e` is `epsilon`, a very small number added to the denominator to prevent it from being zero. - `sqrt` refers to `np.sqrt` or `ops.sqrt`, which is used to take the square root of an expression. After these preparations, I will start explaining each optimizer. --- ## Detailed Explanation of Common Optimizers ### SGD Also known as Stochastic Gradient Descent, it's almost everyone's first encounter with an optimizer. Its principle is simple: randomly select a small batch of data samples, calculate the gradient at each node, and then update the weight of the current node using the learning rate. In Keras 3, its mathematical expression is as follows: ```python w = w - lr * g ``` We can use SGD as a baseline to visually compare the effects of other optimizers: ```python fit_show_model(optimizer=keras.optimizers.SGD(learning_rate=0.01)) ``` ![With SGD, the Loss will converge to 0. in about 500 epochs.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/SGD.png) With SGD, the loss will converge to 0\. in about 500 epochs. Image by Author SGD's biggest drawback is that its learning rate is fixed at every point on the curve, leading to quick weight changes on steep slopes and slow changes on flat ones. This makes it easy to get stuck in local minima (a phenomenon well-documented elsewhere, so I won't repeat it here). How can we avoid getting stuck? Just like helping a car out of the mud, we can give momentum to the weight changes to push it forward, solving this problem. The mathematical expression with `momentum` is shown below: ```python m = momentum * m - lr * g ``` You can see that `momentum` speeds up loss convergence: ```python fit_show_model(optimizer=keras.optimizers.SGD(learning_rate=0.01, momentum=0.9)) ``` ![You can see that momentum speeds up loss convergence.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/SGD_momentum.png) You can see that momentum speeds up loss convergence. Image by Author As mentioned before, since SGD's learning rate is fixed, we need to set a reasonable learning rate. If it's too high, the weight will oscillate back and forth across the valley. If it's too low, the weight will take longer to find the valley, and more likely to get stuck. To address this, besides using `momentum`, we can also predict the next direction of the weight change and add this prediction to the current weight. This method is called the Nesterov method. Its mathematical expression is as follows: ```python m = momentum * m - lr * g w = w + (momentum * m -lr * g) ``` ```python fit_show_model(optimizer=keras.optimizers.SGD(learning_rate=0.01, momentum=0.9, nesterov=True)) ``` ![The Nesterov method allows for smoother convergence of loss.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/SGD_Nesterov.png) The Nesterov method allows for smoother convergence of loss. Image by Author You can see that in Keras 3, using the Nesterov method requires setting the `momentum` parameter first; otherwise, we can't predict the direction of the next weight change. ### Adagrad After discussing algorithms related to stochastic gradient descent, let's talk about optimizers related to adaptive algorithms. The simplest one is Adagrad. The core idea of adaptive algorithms is to dynamically adjust the learning rate as training progresses. Based on this idea, Adagrad adjusts the learning rate by accumulating the sum of the squares of the gradients from historical iterations and dividing the learning rate by this sum. In Keras 3, Adagrad's mathematical expression is as follows: ```python accumulator = accumulator + g**2 w = w - (lr * g) / sqrt(accumulator + e) ``` ```python fit_show_model(optimizer=keras.optimizers.Adagrad(learning_rate=0.01)) ``` ![Adagrad does not converge losses very quickly.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/Adagrad.png) Adagrad does not converge losses very quickly. Image by Author From the chart, we can see that Adagrad does not converge losses very quickly, even slower than the default SGD algorithm, because the learning rate decreases as training progresses. ### RMSprop To address the issue of Adagrad decreasing the learning rate too much, [Hinton proposed an improved algorithm in 2012](http://www.cs.toronto.edu/~tijmen/csc321/slides/lecture%5Fslides%5Flec6.pdf?ref=dataleadsfuture.com). This algorithm doesn't simply add up the squares of the gradients; it assigns a weight to the square of each gradient, giving more weight to recent iterations. The rest of the calculation is similar to Adagrad. ```python v = rho * v + (1 - rho) * g**2 w = w - lr * g / sqrt(v + e) ``` ```python fit_show_model(optimizer=keras.optimizers.RMSprop(learning_rate=0.01), epochs=100) ``` ![RMSprop is much faster than Adagrad and SGD.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/RMSprop.png) RMSprop is much faster than Adagrad and SGD. Image by Author As you can see, although loss convergence isn't very stable, it reaches near zero around 40 epochs, much faster than Adagrad and SGD. In Keras 3, RMSprop also supports setting `momentum` and `centered` parameters. Momentum adds `momentum` to the weight changes, and if the `centered` parameter is set, the optimizer doesn't directly use the accumulated square of the gradients but makes a correction using the moving average of the gradients. The expression is as follows: ```python v = rho * v + (1 - rho) * g**2 average_grad = rho * average_grad + (1 - rho) * g m = momentum * m + (lr * g) / sqrt(v - average_grad**2 + e) w = w - m ``` ```python fit_show_model(optimizer=keras.optimizers.RMSprop(learning_rate=0.01, momentum=0.9, centered=True), epochs=100) ``` ![RMSprop converges losses faster but still has stability issues.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/RMSprop_momentum.png) RMSprop converges losses faster but still has stability issues. Image by Author You can see that using the momentum and centered parameters, RMSprop converges losses faster but still has stability issues. ### Adam Let's talk about the Adam optimizer. Unlike the previous two algorithms, Adam not only uses the accumulated square of the gradients but also the first moment of the gradients. So, Adam has two extra hyperparameters: `beta_1` and `beta_2`. In Keras 3, Adam has evolved further. Now, it adjusts `beta_1` and `beta_2` exponentially based on the current step of the iteration, affecting the size of the learning rate. This evolution makes the Adam optimizer very suitable for time-sensitive scenarios like speech recognition: ```python t = 0 # current iteration local_step = t + 1 beta_1_power = power(beta_1, local_step) beta_2_power = power(beta_2, local_step) alpha = lr * sqrt(1 - beta_2_power) / (1 - beta_1_power) m = m + (1 - beta_1) * (g - m) v = v + (1 - beta_2) * (g**2 - v) w = w - (alpha * m) / (sqrt(v) + e) ``` ```python fit_show_model(optimizer=keras.optimizers.Adam(learning_rate=0.01), epochs=100) ``` ![Adam optimizer converges losses very quickly and smoothly.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/Adam.png) Adam optimizer converges losses very quickly and smoothly. Image by Author From the chart, you can see that the Adam optimizer converges losses very quickly and smoothly, reaching near zero by the 20th epoch. ### AdamW In Keras 3, there is also an optimizer called AdamW, which, as the name suggests, is similar to Adam but adjusts the weights with a constant decay amount through the `weight_decay` parameter. ```python w = w - weight_decay * lr * w ``` ```python fit_show_model(optimizer=keras.optimizers.AdamW(learning_rate=0.01, weight_decay=0.01), epochs=100) ``` ![AdamW is similar to Adam but adjusts the weights with a constant decay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/AdamW.png) AdamW is similar to Adam but adjusts the weights with a constant decay. Image by Author From the source code, it's clear that the AdamW optimizer is actually calling the Adam optimizer and assigning a value to the `weight_decay` parameter. ### Nadam Then there is the Nadam optimizer, which, as the name suggests, is a variant of the Adam optimizer. In Keras 3, it incorporates the idea of Nesterov, not only focusing on the current iteration step but also on the impact of the next step. Then it combines these two effects on the `beta_1` and `beta_2` parameters. So, of all the optimizers, Nadam's algorithm is the most complex: ```bash t = 0 decay = 0.96 local_step = t + 1 next_step = t + 2 u_t = beta_1 * (1.0 - 0.5 * power(decay, local_step)) u_t_1 = beta_2 * (1.0 - 0.5 * power(decay, next_step)) u_product_t = u_product_t * beta_1 * (1.0 - 0.5 * power(decay, local_step)) u_product_t_1 = u_product_t * u_t_1 beta_2_power = power(beta_2, local_step) m = m + (1 - beta_1) * (g - m) v = v + (1 - beta_2) * (g**2 - v) m_hat = u_t_1 * m / (1 - u_product_t_1) + ((1 - u_t) * g) / (1 - u_product_t) v_hat = v / (1 - beta_2_power) w = w - (lr * m_hat) / (sqrt(v_hat) + epsilon) ``` ```python fit_show_model(optimizer=keras.optimizers.Nadam(learning_rate=0.01), epochs=100) ``` ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/Nadam.png) For some simple classification tasks, the improvement from Nadam is not much. Image by Author However, for some simple classification tasks, the improvement from Nadam is not much. ### Lion Finally, let me mention the Lion optimizer, a new implementation proposed by [Chen et al., 2023](http://arxiv.org/abs/2302.06675?ref=dataleadsfuture.com). This optimizer's biggest feature is that it doesn't calculate accumulations, so it uses less memory. It also doesn't calculate the second moment, so it's less complex. It simply calculates a sign based on the current gradient and lets the weight and learning rate change according to this sign: ```python w = w - lr * sign(beta_1 * m + (1 - beta_1) * g) m = beta_2 * m + (1 - beta_2) * g ``` ```python fit_show_model(optimizer=keras.optimizers.Lion(learning_rate=0.001), epochs=100) ``` ![The performance of Lion optimizer isn't as good as the Adam series.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/08/Lion.png) The performance of Lion optimizer isn't as good as the Adam series. Image by Author You can see that the performance of this optimizer isn't as good as the Adam series, probably the price paid for saving resources. --- ## Conclusion As mentioned at the beginning of the article, optimizers, as one of the most critical parts of deep learning, are skills that every practitioner needs to master and apply. Also, with technological advancements, the implementation of optimizers within the latest deep learning frameworks is continuously evolving in terms of computational efficiency and application scenarios. This article introduced the mathematical implementation of several common optimizers in Keras 3. If you're still confused about how to use the optimizers in Keras 3, I suggest starting with Adam. After achieving good results, you can choose a more suitable optimizer based on specific scenarios. What else would you like to know about Keras 3? Feel free to leave comments and discuss. See you next time. ### Keras 3.0 Tutorial: End-to-End Deep Learning Project Guide URL: https://www.dataleadsfuture.com/keras-3-0-tutorial-end-to-end-deep-learning-project-guide/ Last updated: 2026-01-23T06:40:55.000Z ## Introduction Even though I started using Pytorch a while ago, I still miss the concise code style of Keras and the good old days when you could implement a neural network model in just a few lines of code. So, I was thrilled when Keras announced last November that in addition to TensorFlow, it now also supports Pytorch and Jax as backends! However, things weren't perfect: since Keras 3.0 was released not long ago, the related tutorials and documentation hadn't caught up, and I encountered some troubles during the code migration. Luckily, after some effort, I can now smoothly use version 3.0 for various end-to-end model developments. In this article, I'll share some practical experiences with Keras 3.0 to help you avoid some detours. I'll use a typical encoder-decoder recurrent neural network as an example to explain how to complete an end-to-end project from scratch using the subclassing API of Keras 3.0, and discuss details to consider when using Pytorch as the backend. Let's get started. --- ## Framework Installation and Environment Setup ### Framework installation Installing Keras 3.0 (or the latest version) is simple, just follow the [Getting Started](https://keras.io/getting%5Fstarted/?ref=dataleadsfuture.com) documentation on the official website. Before installing Keras, it's recommended to install Pytorch with the corresponding CUDA version first. Either CUDA 11.8 or CUDA 12.1 works, depending on your graphics card driver support. Although Pytorch can be used as a backend, Tensorflow version 2.16.1 is still installed by default during the Keras installation process. This version of Tensorflow is compiled based on CUDA 12.3, so after installing Keras, you might encounter a warning about missing CUDA (see [this issue](https://github.com/tensorflow/tensorflow/issues/63362?ref=dataleadsfuture.com)). ```shell Could not find cuda drivers on your machine, GPU will not be used. ``` Since we are using Pytorch as the backend, my advice is to ignore this warning. Alternatively, you can set a system variable to permanently turn off Tensorflow's logs. ```python import os os.environ['TF_CPP_MIN_LOG_LEVEL'] = '2' ``` ### Environment configuration After installing both Pytorch and Keras, you need to set the environment variable to configure Keras's backend to Pytorch. There are two ways to do this: - Modify the configuration file. - Set an environment variable. First, let's discuss using the configuration file method. Keras's configuration file is located in `~/.keras/keras.json`. If you are using Windows, this file is located in your `/.keras/keras.json`. Of course, you can also change the location of the `.keras` directory by setting the `KERAS_HOME` environment variable. Note, you might not find the `.keras` directory immediately after installing Keras for the first time. At that point, you can execute `import keras` in IPython or Jupyter Notebook to locate the directory. Then, just change the value of the `"backend"` key in the `keras.json` file to `"torch"`. ```json { "floatx": "float32", "epsilon": 1e-07, "backend": "torch", "image_data_format": "channels_last" } ``` If you are in a production system or using cloud environments like Colab, you might not be able to modify the configuration file. In such cases, you can resolve this by setting an environment variable: ```python os.environ["KERAS_BACKEND"] = "torch" ``` Once the Keras backend is configured, you can confirm it with the following code: ```python In: import keras keras.config.backend() Out: 'torch' ``` After the preparations are done, we'll officially start our project practice for today. --- ## Project in Action: An End-to-End Example The fastest way to learn a framework is through real project practice. So now it's time to fulfill my promise. I will guide you through using the subclassing API step by step to implement a neural machine translation (NMT) model and explain some details of using Keras 3.0. ### Theory introduction If you're not familiar with the NMT model, here's a brief introduction: NMT is a type of recurrent neural network model based on an encoder-decoder architecture. In this architecture, there is an embedding layer and an RNN (we use LSTM in this article) layer forming an encoder, and another embedding layer and RNN layer forming a decoder. The original text, after being vectorized, is input into the encoder module. After a series of steps, the final state is input into the decoder module. Additionally, the target text is also input into the decoder module, but before entering the decoder, it is offset by one step forward. Thus, the beginning part of the target text starts with a start-of-sequence (SOS) placeholder. The encoder's input state and the target text's input are processed in the decoder through a series of recurrent calculations, and finally output to a Dense layer, where they are activated to calculate the probabilities of each text vector, compared with the target text's word vectors, and calculate the loss. Therefore, we also add an end-of-sequence (EOS) placeholder at the end of the target text to mark the text's end. The entire architecture is shown in the following diagram: ![The entire architecture is shown in the following diagram. Image by Author](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/05/keras3.0_encoder_decoder-no_logo.webp) The entire architecture is shown in the following diagram. Image by Author Of course, due to the popularity of the Transformer architecture, Keras' [KerasNLP](https://keras.io/keras%5Fnlp/?ref=dataleadsfuture.com) package also offers various pre-trained models like Bert and GPT for completing NLP tasks. However, this article focuses on understanding how to use Keras 3.0, so using a basic RNN network will be enough. ### Modules and flowchart Since this is a production-ready project, we build modules based on Keras 3.0's subclassing API. For a clear understanding of each module and their interactions, I've created the flowchart below: ![The modules and flowchart of this project. Image by Author](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/05/keras3.0_encoder_decoder-flowchart.webp) The modules and flowchart of this project. Image by Author We'll write our code according to the design on the flowchart. ### Importing packages In a Jupyter Notebook environment, I like to import all related packages at the start of the project. This way, if I find something missing in the middle, I just need to add it in one place, instead of searching for the import cell: ```python from pathlib import Path import pickle import keras from keras import layers, utils import numpy as np utils.set_random_seed(42) ``` Here's a little tip: the `utils.set_random_seed` method can set the random seeds for Python, Numpy, and Pytorch all in one line of code, which is super convenient. ### Data preparation Before we start, we need to choose suitable data. Like past encoder-decoder models, we also chose the spa-eng text dataset. > This dataset is provided by contributors to the Tatoeba project and contains 120,000 sentence pairs. It is released under the Creative Commons Attribution 2.0 France license, and you can download the dataset from here. After downloading, let's first check the contents of the spa.txt file: ```txt The rain lasted three days. La lluvia duró tres días. CC-BY 2.0 (France) Attribution: tatoeba.org #27004 (CK) & #431740 (Leono) The refrigerator is closed. El frigorífico está cerrado. CC-BY 2.0 (France) Attribution: tatoeba.org #5152850 (CarpeLanam) & #10211587 (manufrutos) The reports were confusing. Los informes eran confusos. CC-BY 2.0 (France) Attribution: tatoeba.org #2268485 (_undertoad) & #2268486 (cueyayotl) The resemblance is uncanny. La similitud es extraña. CC-BY 2.0 (France) Attribution: tatoeba.org #2691302 (CM) & #5941808 (albrusgher) The resemblance is uncanny. El parecido es asombroso. CC-BY 2.0 (France) Attribution: tatoeba.org #2691302 (CM) & #6026125 (albrusgher) The results seem promising. Los resultados se antojan prometedores. CC-BY 2.0 (France) Attribution: tatoeba.org #8480484 (shekitten) & #8464272 (arh) The rich have many friends. Los ricos tienen muchos amigos. CC-BY 2.0 (France) Attribution: tatoeba.org #1579047 (sam_m) & #1457378 (marcelostockle) ``` As you can see, the content includes at least three columns, with the first column being the original text and the second column being the target text, separated by tabs. Since the file isn't large, we can directly use numpy's `genfromtxt` method to read this dataset. ```python text_file = Path("./temp/eng-spanish/spa-eng/spa.txt") pairs = np.genfromtxt(text_file, delimiter="\t", dtype=str, usecols=(0, 1), encoding="utf-8", autostrip=True, converters={1: lambda x: x.replace("¡", "").replace("¿", "")}) np.random.shuffle(pairs) sentence_en, sentence_es = pairs[:, 0], pairs[:, 1] ``` Next, let's check the processing results: ```python In: print(f"{sentence_en[0]} => {sentence_es[0]}") Out: I'm really sorry. => Realmente lo siento. ``` OK, no problems. ### Data preprocessing Next, we need to preprocess the text content to convert it into word vector data. First, we define some constants: ```python class Configure: VOCAB_SIZE: int = 1000 MAX_LENGTH: int = 50 SOS: str = 'startofseq' EOS: str = 'endofseq' ``` Then, we start our data processing pipeline. Note that in Keras 3.0, although you have chosen Pytorch as the backend, the `TextVectorization` Layer is still implemented based on TensorFlow. Therefore, you cannot use `TextVectorization` as a layer in the Keras Model but must use it separately in the preprocessing pipeline. This leads to a problem: when we migrate the trained model to the production system for inference tasks, without the TextVectorization vocabulary, we cannot perform vectorization. So, we need to persist the vocabulary and reuse it, but there are some issues with the persistence of Keras 3.0's `TextVectorization`, which I will discuss later. I will use a `TextPreprocessor` module to perform the vectorization. Here is the specific code: ```python class TextPreprocessor: def __init__(self, en_config = None, es_config = None): if en_config is None: self.text_vec_layer_en = layers.TextVectorization( Configure.VOCAB_SIZE, output_sequence_length=Configure.MAX_LENGTH ) else: self.text_vec_layer_en = layers.TextVectorization.from_config(en_config) if es_config is None: self.text_vec_layer_es = layers.TextVectorization( Configure.VOCAB_SIZE, output_sequence_length=Configure.MAX_LENGTH ) else: self.text_vec_layer_es= layers.TextVectorization.from_config(es_config) self.adapted = False self.sos = Configure.SOS self.eos = Configure.EOS def adapt(self, en_sentences: list[str], es_sentences: list[str]) -> None: self.text_vec_layer_en.adapt(en_sentences) self.text_vec_layer_es.adapt([f"{self.sos} {s} {self.eos}" for s in es_sentences]) self.adapted = True def en_vocabulary(self): return self.text_vec_layer_en.get_vocabulary() def es_vocabulary(self): return self.text_vec_layer_es.get_vocabulary() def vectorize_en(self, en_sentences: list[str]): return self.text_vec_layer_en(en_sentences) def vectorize_es(self, es_sentences: list[str]): return self.text_vec_layer_es(es_sentences) @classmethod def from_config(cls, config): return cls(**config) def get_config(self): en_config = self.text_vec_layer_en.get_config() en_config['vocabulary'] = self.en_vocabulary() es_config = self.text_vec_layer_es.get_config() es_config['vocabulary'] = self.es_vocabulary() return {'en_config': en_config, 'es_config': es_config} def save(self, filepath: str): if not self.adapted: raise RuntimeError("Layer hasn't been adapted yet.") if filepath is None: raise ValueError("A file path needs to be defined.") if not filepath.endswith('.pkl'): raise ValueError("The file path needs to end in .pkl.") pickle.dump({ 'config': self.get_config() }, open(filepath, 'wb')) @classmethod def load(cls, filepath: str): conf = pickle.load(open(filepath, 'rb')) instance = cls(**conf['config']) return instance ``` Let me explain what this module does: - Since we need to vectorize both the original text and the target text, this module includes two `TextVectorization` Layers. - After adapting, this module will hold the vocabularies for both the original and target texts. This way, when deploying to the production system, the `TextVectorization` won't need to adapt again. - The module uses the pickle module to enable persistence. You can use the `get_config` method to get the configuration of the two `TextVectorization` Layers and save it. You can also use `from_config` to initialize the module's instance from the saved configuration directly. - However, when I used the `get_config` method, the vocabulary wasn't retrieved (currently, I'm using Keras version 3.3, and I'm not sure if this is a bug), so I had to use the `get_vocabulary` method to get the vocabulary separately. Let's adapt the text and save the vocabulary: ```python text_preprocessor = TextPreprocessor() text_preprocessor.adapt(sentence_en, sentence_es) text_preprocessor.save('./data/text_preprocessor.pkl') ``` Check the vocabularies for both languages: ```python In: text_preprocessor.en_vocabulary()[:10] Out: ['', '[UNK]', 'i', 'the', 'to', 'you', 'tom', 'a', 'is', 'he'] In: text_preprocessor.es_vocabulary()[:10] Out: ['', '[UNK]', 'startofseq', 'endofseq', 'de', 'que', 'no', 'tom', 'a', 'la'] ``` No problem at all. Once the TextPreprocessor module is ready, we can start splitting the training and validation sets and begin the vectorization work. Since the target text also serves as input for the decoder module, we have two additional feature sets: `X_train_dec` and `X_valid_dec`: ```python X_train = text_preprocessor.vectorize_en(sentence_en[:100_000]) X_valid = text_preprocessor.vectorize_en(sentence_en[100_000:]) X_train_dec = text_preprocessor.vectorize_es([f"{Configure.SOS} {s}" for s in sentence_es[:100_000]]) X_valid_dec = text_preprocessor.vectorize_es([f"{Configure.SOS} {s}" for s in sentence_es[100_000:]]) y_train = text_preprocessor.vectorize_es([f"{s} {Configure.EOS}" for s in sentence_es[:100_000]]) y_valid = text_preprocessor.vectorize_es([f"{s} {Configure.EOS}" for s in sentence_es[100_000:]]) ``` ### Implementing the encoder-decoder model As depicted in the architecture diagram earlier, the entire model is divided into encoder and decoder parts. So, we implement two custom subclasses based on `keras.layers.Layer` for each part. It's important to implement the `__init__`, `call`, and `get_config` methods for each custom Layer. - The `__init__` method initializes the Layer's member variables, weights, and sub-layers. - The `call` method works similarly to Keras's Functional API, accepting inputs as parameters and returning the Layer's output after processing. - The `get_config` method is used to retrieve the configuration of the Layer when saving the model. Encoder Layer: ```python @keras.saving.register_keras_serializable() class Encoder(keras.layers.Layer): def __init__(self, embed_size: int = 128, **kwargs): super().__init__(**kwargs) self.embed_size = embed_size self.encoder_embedding_layer = layers.Embedding(input_dim=Configure.VOCAB_SIZE, output_dim=self.embed_size, mask_zero=True) self.encoder = layers.LSTM(512, return_state=True) def call(self, inputs): encoder_embeddings = self.encoder_embedding_layer(inputs) encoder_outputs, *encoder_state = self.encoder(encoder_embeddings) return encoder_outputs, encoder_state def get_config(self): config = {"embed_size": self.embed_size} base_config = super().get_config() return config | base_config ``` In the Encoder, we set the `return_state` parameter of `LSTM` to True. This allows the final state of the `LSTM` to be returned as output for the Decoder Layer to use. Decoder Layer: ```python @keras.saving.register_keras_serializable() class Decoder(keras.layers.Layer): def __init__(self, embed_size: int = 128, **kwargs): super().__init__(**kwargs) self.embed_size = embed_size self.decoder_embedding_layer = layers.Embedding(input_dim=Configure.VOCAB_SIZE, output_dim=self.embed_size, mask_zero=True) self.decoder = layers.LSTM(512, return_sequences=True) def call(self, inputs, initial_state=None): decoder_embeddings = self.decoder_embedding_layer(inputs) decoder_outputs = self.decoder(decoder_embeddings, initial_state=initial_state) return decoder_outputs def get_config(self): config = {"embed_size": self.embed_size} base_config = super().get_config() return config | base_config ``` In the Decoder, besides receiving data input, the call method also accepts the Encoder's input through the `initial_state` function and returns the module's output. We also implement a custom Model, which needs to implement the `__init__`, `call`, and `get_config` methods, similar to `keras.layers.Layer`. ```python @keras.saving.register_keras_serializable() class NMTModel(keras.models.Model): embed_size: int = 128 def __init__(self, **kwargs): super().__init__(**kwargs) self.encoder = Encoder(self.embed_size) self.decoder = Decoder(self.embed_size) self.out = layers.Dense(Configure.VOCAB_SIZE, activation='softmax') def call(self, inputs): encoder_inputs, decoder_inputs = inputs encoder_outputs, encoder_state = self.encoder(encoder_inputs) decoder_outputs = self.decoder(decoder_inputs, initial_state=encoder_state) out_proba = self.out(decoder_outputs) return out_proba def get_config(self): base_config = super().get_config() return base_config ``` - In the Model, we initialize a `Dense` layer to convert the Decoder's output into results for the word vectors. - The `call` method takes two inputs, which can be easily distinguished through unpacking. - Both Layer and Model need to have the `@keras.saving.register_keras_serializable()` decorator to ensure correct serialization when saving the model. ### Model training After defining the model, we proceed to the training phase: ```python nmt_model = NMTModel() nmt_model.compile(loss='sparse_categorical_crossentropy', optimizer='nadam', metrics=['accuracy']) checkpoint = keras.callbacks.ModelCheckpoint( './data/nmt_model.keras', monitor='val_accuracy', save_best_only=True ) nmt_model.fit((X_train, X_train_dec), y_train, epochs=1, validation_data=((X_valid, X_valid_dec), y_valid), batch_size=128, callbacks=[checkpoint]) ``` In this part of the code: - We first call the compile method to compile the model instance, defining components such as `loss`, `optimizer`, and `metrics`. - We set up a `ModelCheckpoint` callback to save the model with the best `val_accuracy` after training. - We use the `fit` method, passing `X_train` and `X_train_dec` as a tuple to the `x` parameter, and handle `validation_data` similarly. - This is just a demo, so I set `epochs` to 1\. You can adjust the values of `epochs` and `batch_size` as needed. - Keras 3.0 also supports Pytorch's `DataLoader`, or you can implement a backend-agnostic preprocessing pipeline based on [keras.utils.PyDataset](https://keras.io/api/utils/python%5Futils/?ref=dataleadsfuture.com#pydataset-class). I can explain how to use these in my next article. After training is complete, the model should be saved. ### Inference task After training, you can deploy the corresponding code modules, along with the saved vocabulary and model, to the production system for inference tasks. Since the model's `Dense` layer outputs the probability of each word vector in the vocabulary, you need to merge each inferred word with the previous results and re-input them with the original text to predict the next word: ```python preprocessor = TextPreprocessor.load('./data/text_preprocessor.pkl') nmt_model = keras.saving.load_model('./data/nmt_model.keras') def translate(sentence_en): translation = "" for word_index in range(50): X = preprocessor.vectorize_en([sentence_en]) X_dec = preprocessor.vectorize_es([Configure.SOS + " " + translation]) y_proba = nmt_model.predict((X, X_dec), verbose=0)[0, word_index] predicted_word_id = np.argmax(y_proba) predicted_word = preprocessor.es_vocabulary()[predicted_word_id] if predicted_word == Configure.EOS: break translation = translation + " " + predicted_word return translation.strip() ``` Let's write a simple method to test the results: ```python In: translate("It was pretty cool.") Out: 'era bastante [UNK]' ``` Although it's not very accurate, the goal of this article is to learn how to use the Keras 3.0 subclassing API, so you still have plenty of room to optimize this model, right? --- ## Conclusion The release of Keras 3.0 allows us to implement models efficiently using Keras's concise API while using Pytorch or Jax as backends. However, since the version was released recently, the accompanying documentation is not yet complete, so you might encounter some difficulties in trying new versions. This article through an end-to-end practical example, explains the environment setup and basic development process of Keras 3.0, helping you get started quickly. Unfortunately, the Keras 3.0 project is still in its early stages and cannot completely break away from the dependence on TensorFlow, as well as some of TensorFlow's inexplicable issues. But I am still optimistic about this version. I believe that as time goes on and support for multiple backends improves, Keras will be revitalized, helping to make deep learning technology more accessible and reducing the learning curve for deep learning. What else would you like to know about Keras 3.0? Feel free to leave a comment and discuss. ### Scikit-learn Visualization Guide: Making Models Speak URL: https://www.dataleadsfuture.com/scikit-learn-visualization-guide-making-models-speak/ Last updated: 2025-03-26T02:15:56.000Z ## Introduction In the journey of machine learning, explaining models with visualization is as important as training them. A good chart can show us what a model is doing in an easy-to-understand way. Here's an example: ![Decision boundaries of two different generalization performances.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/DecisionBoundarys.jpg) Decision boundaries of two different generalization performances. Image by Author This graph makes it clear that for the same dataset, the model on the right is better at generalizing. Most machine learning books prefer to use raw Matplotlib code for visualization, which leads to issues: 1. You have to learn a lot about drawing with Matplotlib. 2. Plotting code fills up your notebook, making it hard to read. 3. Sometimes you need third-party libraries, which isn't ideal in business settings. Good news! Scikit-learn now offers `Display` classes that let us use methods like `from_estimator` and `from_predictions` to make drawing graphs for different situations much easier. Curious? Let me show you these cool APIs. --- ## Scikit-learn Display API Introduction ### Use utils.discovery.all\_displays to find available APIs Scikit-learn (sklearn) always adds `Display` APIs in new releases, so it's key to know what's available in your version. Sklearn's [utils.discovery.all\_displays](https://scikit-learn.org/stable/modules/generated/sklearn.utils.discovery.all%5Fdisplays.html?ref=dataleadsfuture.com#sklearn.utils.discovery.all%5Fdisplays) lets you see which classes you can use. ```Python from sklearn.utils.discovery import all_displays displays = all_displays() displays ``` For example, in my Scikit-learn 1.4.0, these classes are available: ```Python [('CalibrationDisplay', sklearn.calibration.CalibrationDisplay), ('ConfusionMatrixDisplay', sklearn.metrics._plot.confusion_matrix.ConfusionMatrixDisplay), ('DecisionBoundaryDisplay', sklearn.inspection._plot.decision_boundary.DecisionBoundaryDisplay), ('DetCurveDisplay', sklearn.metrics._plot.det_curve.DetCurveDisplay), ('LearningCurveDisplay', sklearn.model_selection._plot.LearningCurveDisplay), ('PartialDependenceDisplay', sklearn.inspection._plot.partial_dependence.PartialDependenceDisplay), ('PrecisionRecallDisplay', sklearn.metrics._plot.precision_recall_curve.PrecisionRecallDisplay), ('PredictionErrorDisplay', sklearn.metrics._plot.regression.PredictionErrorDisplay), ('RocCurveDisplay', sklearn.metrics._plot.roc_curve.RocCurveDisplay), ('ValidationCurveDisplay', sklearn.model_selection._plot.ValidationCurveDisplay)] ``` ### Using inspection.DecisionBoundaryDisplay for decision boundaries Since we mentioned it, let's start with decision boundaries. If you use Matplotlib to draw them, it's a hassle: - Use `np.linspace` to set coordinate ranges; - Use `plt.meshgrid` to calculate the grid; - Use `plt.contourf` to draw the decision boundary fill; - Then use `plt.scatter` to plot data points. Now, with [inspection.DecisionBoundaryDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.inspection.DecisionBoundaryDisplay.html?ref=dataleadsfuture.com#sklearn-inspection-decisionboundarydisplay), you can simplify this process: ```Python from sklearn.inspection import DecisionBoundaryDisplay from sklearn.datasets import load_iris from sklearn.svm import SVC from sklearn.pipeline import make_pipeline from sklearn.preprocessing import StandardScaler import matplotlib.pyplot as plt iris = load_iris(as_frame=True) X = iris.data[['petal length (cm)', 'petal width (cm)']] y = iris.target svc_clf = make_pipeline(StandardScaler(), SVC(kernel='linear', C=1)) svc_clf.fit(X, y) display = DecisionBoundaryDisplay.from_estimator(svc_clf, X, grid_resolution=1000, xlabel="Petal length (cm)", ylabel="Petal width (cm)") plt.scatter(X.iloc[:, 0], X.iloc[:, 1], c=y, edgecolors='w') plt.title("Decision Boundary") plt.show() ``` See the final effect in the figure: ![Use DecisionBoundaryDisplay to draw a triple classification model.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/DecisionBoundaryDisplay.webp) Use DecisionBoundaryDisplay to draw a triple classification model. Image by Author Remember, `Display` can only draw 2D, so make sure your data has only two features or reduced dimensions. ### Using calibration.CalibrationDisplay for probability calibration When your classification model outputs predictions as probabilities, you can use a probability calibration curve to show how well the predicted probabilities match the actual probabilities. For a perfect classification model, its calibration curve will follow a 45-degree line, meaning the predicted probabilities are consistent with the actual probabilities. Note that [CalibrationDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.calibration.CalibrationDisplay.html?ref=dataleadsfuture.com#sklearn.calibration.CalibrationDisplay) uses the model's `predict_proba`. If you use a support vector machine, set `probability` to True: ```Python from sklearn.calibration import CalibrationDisplay from sklearn.model_selection import train_test_split from sklearn.datasets import make_classification from sklearn.ensemble import HistGradientBoostingClassifier X, y = make_classification(n_samples=1000, n_classes=2, n_features=5, random_state=42) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42) proba_clf = make_pipeline(StandardScaler(), SVC(kernel="rbf", gamma="auto", C=10, probability=True)) proba_clf.fit(X_train, y_train) CalibrationDisplay.from_estimator(proba_clf, X_test, y_test) hist_clf = HistGradientBoostingClassifier() hist_clf.fit(X_train, y_train) ax = plt.gca() CalibrationDisplay.from_estimator(hist_clf, X_test, y_test, ax=ax) plt.show() ``` ![Charts drawn by CalibrationDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/calibration.webp) Charts drawn by CalibrationDisplay. Image by Author ### Using metrics.ConfusionMatrixDisplay for confusion matrices When assessing classification models and dealing with imbalanced data, we look at precision and recall. These break down into TP, FP, TN, and FN – a confusion matrix. To draw one, use [metrics.ConfusionMatrixDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.ConfusionMatrixDisplay.html?ref=dataleadsfuture.com#sklearn-metrics-confusionmatrixdisplay). It's well-known, so I'll skip the details. ```Python from sklearn.datasets import fetch_openml from sklearn.ensemble import RandomForestClassifier from sklearn.metrics import ConfusionMatrixDisplay digits = fetch_openml('mnist_784', version=1) X, y = digits.data, digits.target rf_clf = RandomForestClassifier(max_depth=5, random_state=42) rf_clf.fit(X, y) ConfusionMatrixDisplay.from_estimator(rf_clf, X, y) plt.show() ``` ![Charts drawn with ConfusionMatrixDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/ConfusionMatrixDisplay.webp) Charts drawn with ConfusionMatrixDisplay. Image by Author ### metrics.RocCurveDisplay and metrics.DetCurveDisplay These two are together because they're often used to evaluate side by side. [RocCurveDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.RocCurveDisplay.html?ref=dataleadsfuture.com#sklearn.metrics.RocCurveDisplay) compares TPR and FPR for the model. For binary classification, you want low FPR and high TPR, so the upper left corner is best. The Roc curve bends towards this corner. Because the Roc curve stays near the upper left, leaving the lower right empty, it's hard to see model differences. So, we also use [DetCurveDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.DetCurveDisplay.html?ref=dataleadsfuture.com#sklearn.metrics.DetCurveDisplay) to draw a Det curve with FNR and FPR. It uses more space, making it clearer than the Roc curve. The perfect point for a Det curve is the lower left corner. ```Python from sklearn.metrics import RocCurveDisplay from sklearn.metrics import DetCurveDisplay X, y = make_classification(n_samples=10_000, n_features=5, n_classes=2, n_informative=2) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, random_state=42, stratify=y) classifiers = { "SVC": make_pipeline(StandardScaler(), SVC(kernel="linear", C=0.1, random_state=42)), "Random Forest": RandomForestClassifier(max_depth=5, random_state=42) } fig, [ax_roc, ax_det] = plt.subplots(1, 2, figsize=(10, 4)) for name, clf in classifiers.items(): clf.fit(X_train, y_train) RocCurveDisplay.from_estimator(clf, X_test, y_test, ax=ax_roc, name=name) DetCurveDisplay.from_estimator(clf, X_test, y_test, ax=ax_det, name=name) ``` ![Comparison Chart of RocCurveDisplay and DetCurveDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/roc-curve.webp) Comparison Chart of RocCurveDisplay and DetCurveDisplay. Image by Author ### Using metrics.PrecisionRecallDisplay to adjust thresholds With imbalanced data, you might want to shift recall and precision. - For email fraud, you want high precision. - For disease screening, you want high recall to catch more cases. You can adjust the threshold, but what's the right amount? Here, [metrics.PrecisionRecallDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.PrecisionRecallDisplay.html?ref=dataleadsfuture.com#sklearn-metrics-precisionrecalldisplay) can help. ```Python from xgboost import XGBClassifier from sklearn.datasets import load_wine from sklearn.metrics import PrecisionRecallDisplay wine = load_wine() X, y = wine.data[wine.target<=1], wine.target[wine.target<=1] X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.3, stratify=y, random_state=42) xgb_clf = XGBClassifier() xgb_clf.fit(X_train, y_train) PrecisionRecallDisplay.from_estimator(xgb_clf, X_test, y_test) plt.show() ``` ![Charting xgboost model evaluation using PrecisionRecallDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/precision-recall.webp) Charting xgboost model evaluation using PrecisionRecallDisplay. Image by Author This shows that models following Scikit-learn's design can be drawn, like `xgboost` here. Handy, right? ### Using metrics.PredictionErrorDisplay for regression models We've talked about classification, now let's talk about regression. Scikit-learn's [metrics.PredictionErrorDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.PredictionErrorDisplay.html?ref=dataleadsfuture.com#sklearn-metrics-predictionerrordisplay) helps assess regression models. ```Python from sklearn.svm import SVR from sklearn.metrics import PredictionErrorDisplay rng = np.random.default_rng(42) X = rng.random(size=(200, 2)) * 10 y = X[:, 0]**2 + 5 * X[:, 1] + 10 + rng.normal(loc=0.0, scale=0.1, size=(200,)) reg = make_pipeline(StandardScaler(), SVR(kernel='linear', C=10)) reg.fit(X, y) fig, axes = plt.subplots(1, 2, figsize=(8, 4)) PredictionErrorDisplay.from_estimator(reg, X, y, ax=axes[0], kind="actual_vs_predicted") PredictionErrorDisplay.from_estimator(reg, X, y, ax=axes[1], kind="residual_vs_predicted") plt.show() ``` ![Two charts were drawn by PredictionErrorDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/prediction-error-1.webp) Two charts were drawn by PredictionErrorDisplay. Image by Author As shown, it can draw two kinds of graphs. The left shows predicted vs. actual values – good for linear regression. However, not all data is perfectly linear. For that, use the right graph. It compares real vs. predicted differences, a residuals plot. This plot's banana shape suggests our data might not fit linear regression. Switching from a linear to an `rbf` kernel can help. ```Python reg = make_pipeline(StandardScaler(), SVR(kernel='rbf', C=10)) ``` ![A visual demonstration of the improved model performance. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/prediction-error-2.webp) A visual demonstration of the improved model performance. Image by Author See, with `rbf`, the residual plot looks better. ### Using model\_selection.LearningCurveDisplay for learning curves After assessing performance, let's look at optimization with [LearningCurveDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.model%5Fselection.LearningCurveDisplay.html?ref=dataleadsfuture.com#sklearn.model%5Fselection.LearningCurveDisplay). First up, learning curves – how well the model generalizes with different training and testing data, and if it suffers from variance or bias. As shown below, we compare a `DecisionTreeClassifier` and a `GradientBoostingClassifier` to see how they do as training data changes. ```Python from sklearn.tree import DecisionTreeClassifier from sklearn.ensemble import GradientBoostingClassifier from sklearn.model_selection import LearningCurveDisplay X, y = make_classification(n_samples=1000, n_classes=2, n_features=10, n_informative=2, n_redundant=0, n_repeated=0) tree_clf = DecisionTreeClassifier(max_depth=3, random_state=42) gb_clf = GradientBoostingClassifier(n_estimators=50, max_depth=3, tol=1e-3) train_sizes = np.linspace(0.4, 1.0, 10) fig, axes = plt.subplots(1, 2, figsize=(10, 4)) LearningCurveDisplay.from_estimator(tree_clf, X, y, train_sizes=train_sizes, ax=axes[0], scoring='accuracy') axes[0].set_title('DecisionTreeClassifier') LearningCurveDisplay.from_estimator(gb_clf, X, y, train_sizes=train_sizes, ax=axes[1], scoring='accuracy') axes[1].set_title('GradientBoostingClassifier') plt.show() ``` ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/LearningCurveDisplay.webp) Comparison of the learning curve of two different models. Image by Author The graph shows that although the tree-based GradientBoostingClassifier maintains good accuracy on the training data, its generalization capability on test data does not have a significant advantage over the DecisionTreeClassifier. ### Using model\_selection.ValidationCurveDisplay for visualizing parameter tuning So, for models that don't generalize well, you might try adjusting the model's regularization parameters to tweak its performance. The traditional approach is to use tools like `GridSearchCV` or `Optuna` to tune the model, but these methods only give you the overall best-performing model and the tuning process is not very intuitive. For scenarios where you want to adjust a specific parameter to test its effect on the model, I recommend using [model\_selection.ValidationCurveDisplay](https://scikit-learn.org/stable/modules/generated/sklearn.model%5Fselection.ValidationCurveDisplay.html?ref=dataleadsfuture.com#sklearn.model%5Fselection.ValidationCurveDisplay) to visualize how the model performs as the parameter changes. ```Python from sklearn.model_selection import ValidationCurveDisplay from sklearn.linear_model import LogisticRegression param_name, param_range = "C", np.logspace(-8, 3, 10) lr_clf = LogisticRegression() ValidationCurveDisplay.from_estimator(lr_clf, X, y, param_name=param_name, param_range=param_range, scoring='f1_weighted', cv=5, n_jobs=-1) plt.show() ``` ![Fine-tuning of model parameters plotted with ValidationCurveDisplay.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/03/validation-curve.jpg) Fine-tuning of model parameters plotted with ValidationCurveDisplay. Image by Author --- ## Some regrets After trying out all these Displays, I must admit some regrets: - The biggest one is that most of these APIs lack detailed tutorials, which is probably why they're not well-known compared to Scikit-learn's thorough documentation. - These APIs are scattered across various packages, making it hard to reference them from a single place. - The code is still pretty basic. You often need to pair it with Matplotlib's APIs to get the job done. A typical example is `DecisionBoundaryDisplay`, where after plotting the decision boundary, you still need Matplotlib to plot the data distribution. - They're hard to extend. Besides a few methods validating parameters, it's tough to simplify my model visualization process with tools or methods; I end up rewriting a lot. I hope these APIs get more attention, and as versions upgrade, visualization APIs become even easier to use. --- ## Conclusion In the journey of machine learning, explaining models with visualization is as important as training them. This article introduced various plotting APIs in the current version of scikit-learn. With these APIs, you can simplify some Matplotlib code, ease your learning curve, and streamline your model evaluation process. Due to length, I didn't expand on each API. If interested, you can check the [official documentation](https://scikit-learn.org/stable/visualizations.html?ref=dataleadsfuture.com) for more details. Now it's your turn. What are your expectations for visualizing machine learning methods? Feel free to leave a comment and discuss. ### Visualizing What Batch Normalization Is and Its Advantages URL: https://www.dataleadsfuture.com/visualizing-what-batch-normalization-is-and-its-advantages/ Last updated: 2025-03-26T02:15:28.000Z ## Introduction Have you, when conducting deep learning projects, ever encountered a situation where the more layers your neural network has, the slower the training becomes? If your answer is YES, then congratulations, it's time for you to consider using batch normalization now. --- ## What is Batch Normalization? As the name suggests, batch normalization is a technique where batched training data, after activation in the current layer and before moving to the next layer, is standardized. Here's how it works: 1. *The entire dataset is randomly divided into N batches without replacement, each with a mini\_batch size, for the training.* 2. *For the i-th batch, standardize the data distribution within the batch using the formula: `(Xi - Xmean) / Xstd`.* 3. *Scale and shift the standardized data with `γXi + β` to allow the neural network to undo the effects of standardization if needed.* The steps seem simple, don't they? So, what are the advantages of batch normalization? --- ## Advantages of Batch Normalization ### Speeds up model convergence Neural networks commonly adjust parameters using gradient descent. If the cost function is smooth and has only one lowest point, the parameters will converge quickly along the gradient. But if there's a significant variance in the data distribution across nodes, the cost function becomes less like a pit bottom and more like a valley, making the convergence of the gradient exceptionally slow. Confused? No worries, let's explain this situation with a visual: First, prepare a virtual dataset with only two features, where the distribution of features is vastly different, along with a target function: ```Python rng = np.random.default_rng(42) A = rng.uniform(1, 10, 100) B = rng.uniform(1, 200, 100) y = 2*A + 3*B + rng.normal(size=100) * 0.1 # with a little bias ``` Then, with the help of GPT, we use matplot3d to visualize the gradient descent situation before data standardization: ![Visualization of cost functions without standardization of data.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/un_normalized_p.webp) Visualization of cost functions without standardization of data. Image by Author Notice anything? Because one feature's span is too large, the function's gradient is stretched long in the direction of this feature, creating a valley. Now, for the gradient to reach the bottom of the cost function, it has to go through many more iterations. But what if we standardize the two features first? ```Python def normalize(X): mean = np.mean(X) std = np.std(X) return (X - mean)/std A = normalize(A) B = normalize(B) ``` Let's look at the cost function after data standardization: ![Visualization of standardized cost functions for data. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/normalized_p.webp) Visualization of standardized cost functions for data. Image by Author Clearly, the function turns into the shape of a bowl. The gradient simply needs to descend along the slope to reach the bottom. Isn't that much faster? ### Slows down the problem of gradient vanishing The graph we just used has already demonstrated this advantage, but let's take a closer look. Remember this function? ![Visualization of sigmoid function.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/sigmoid.png) Visualization of sigmoid function. Image by Author Yes, that's the sigmoid function, which many neural networks use as an activation function. Looking closely at the sigmoid function, we find that the slope is steepest between -2 and 2. ![The slope of the sigmoid function is steepest between -2 and 2.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/sigmoid_1.webp) The slope of the sigmoid function is steepest between -2 and 2\. Image by Author If we reduce the standardized data to a straight line, we'll find that these data are distributed exactly within the steepest slope of the sigmoid. At this point, we can consider the gradient to be descending the fastest. ![The normalized data will be distributed in the steepest interval of the sigmoid function. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/sigmoid_2.webp) The normalized data will be distributed in the steepest interval of the sigmoid function. Image by Author However, as the network goes deeper, the activated data will drift layer by layer (Internal Covariate Shift), and a large amount of data will be distributed away from the zero point, where the slope gradually flattens. ![The distribution of data is progressively shifted within the neural network.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/sigmoid_3.webp) The distribution of data is progressively shifted within the neural network. Image by Author At this point, the gradient descent becomes slower and slower, which is why with more neural network layers, the convergence becomes slower. If we standardize the data of the mini\_batch again after each layer's activation, the data for the current layer will return to the steeper slope area, and the problem of gradient vanishing can be greatly alleviated. ![The renormalized data return to the region with the steepest slope.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/sigmoid_4.webp) The renormalized data return to the region with the steepest slope. Image by Author ### Has a regularizing effect If we don't batch the training and standardize the entire dataset directly, the data distribution would look like the following: ![Distribution after normalizing the entire data set. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/uniform_distribution.png) Distribution after normalizing the entire data set. Image by Author However since we divide the data into several batches and standardize the data according to the distribution within each batch, the data distribution will be slightly different. ![Distribution of data sets after normalization by batch.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/02/separated_distribution.png) Distribution of data sets after normalization by batch. Image by Author You can see that the data distribution has some minor noise, similar to the noise introduced by Dropout, thus providing a certain level of regularization for the neural network. --- ## Conclusion Batch normalization is a technique that standardizes the data from different batches to accelerate the training of neural networks. It has the following advantages: - Speeds up model convergence. - Slows down the problem of gradient vanishing. - Has a regularizing effect. Have you learned something new? Now it's your turn. What other techniques do you know that optimize neural network performance? Feel free to leave a comment and discuss. ### Ensuring Correct Use of Transformers in Scikit-learn Pipeline URL: https://www.dataleadsfuture.com/ensuring-correct-use-of-transformers-in-scikit-learn-pipeline/ Last updated: 2025-03-26T02:15:02.000Z This article will explain how to use [Pipeline](https://scikit-learn.org/stable/modules/compose.html?ref=dataleadsfuture.com) and [Transformers](https://scikit-learn.org/stable/data%5Ftransforms.html?ref=dataleadsfuture.com) correctly in Scikit-Learn (sklearn) projects to speed up and reuse our model training process. This piece complements and clarifies the official documentation on Pipeline examples and some common misunderstandings. I hope that after reading this, you'll be able to use the Pipeline, an excellent design, to better complete your machine learning tasks. --- ## Introduction There's a famous dish in Chinese restaurants around the world called "General Tso's Chicken," and I wonder if you've tried it. ![General Tso's Chicken.A model for standardizing the cooking process.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2024/01/General-Tso-s-Chicken.webp) General Tso's Chicken.A model for standardizing the cooking process. Photo Credit: Created by Author, Canva. One characteristic of "General Tso's Chicken" is that each piece of chicken is processed by the chef to be the same size. This ensures that: 1. All pieces are marinated for the same amount of time. 2. During cooking, each piece of chicken reaches the same level of doneness. 3. When using chopsticks, the uniform size makes it easier to pick up the pieces. This preprocessing includes washing, cutting, and marinating the ingredients. If the chicken pieces are cut larger than usual, the flavor can change significantly even if stir-fried for the same amount of time. So, when preparing to open a restaurant, we must consider standardizing these processes and recipes to ensure that each plate of "General Tso's Chicken" has a consistent taste and texture. This is how restaurants thrive. Back in the world of machine learning, Scikit-Learn also provides such standardized processes called Pipeline. They solidify the data preprocessing and model training process into a standardized workflow, making machine learning projects easier to maintain and reuse. In this article, we'll explore how to use Transformers correctly within Scikit-Learn's Pipeline, ensuring that our data is as perfectly prepared as the ingredients for a fine meal. --- ## Why Use Transformers ### What are Transformers In Scikit-Learn, Transformers mainly fall into two categories: data scaling and feature dimensionality reduction. Take, for example, a set of housing data, which includes features like location, area, and number of bedrooms. If you don't standardize these features to the same scale, the model might overlook the significant impact of location (usually categorical data) due to minor fluctuations in the area (usually a larger numerical value). It's like overpowering the delicate taste of herbs with too much pepper. ### Using Transformers correctly Typically, data scaling is done using standardization, formulated as: ![Formula for train_data's standardization.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image.png) Formula for train\_data's standardization. Image by Author Where `train_mean` and `train_std` are variables extracted from the train data. In Scikit-Learn, train data and test data are both obtained from the original dataset using the [train\_test\_split](https://scikit-learn.org/stable/modules/generated/sklearn.model%5Fselection.train%5Ftest%5Fsplit.html?ref=dataleadsfuture.com#sklearn.model%5Fselection.train%5Ftest%5Fsplit) method. When scaling test data, the same `train_mean` and `train_std` are used: ![The same train_mean and train_std variables are used when scaling test data.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-1.png) The same train\_mean and train\_std variables are used when scaling test data. Image by Author Here arises the question: why use train data to generate these variables? Let's look at a simple dataset where the train data is: ![A simple dataset of train data.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-2.png) A simple dataset of train data. Image by Author After standardization, the train data becomes: ![The simple dataset after scaling.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-3.png) The simple dataset after scaling. Image by Author Clearly, after scaling, features greater than 0 have a label of 1, which means features greater than 10 before scaling have a label of 1. Now let's look at the test data: ![Test data that has not yet been classified.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-4.png) Test data that has not yet been classified. Image by Author If we use `test_mean` and `test_std` generated from the test data distribution without considering the train data, the results become: ![Demonstration of errors in using test data to generate variables.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-5.png) Demonstration of errors in using test data to generate variables. Image by Author Obviously, this prediction result does not make sense. But suppose we use `train_mean` and `train_std` to process the data and combine it with the model prediction; let's see what happens: ![Using the variables of train data, we obtained the correct results.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-6.png) Using the variables of train data, we obtained the correct results. Image by Author As we can see, only by preprocessing the data with variables generated through train data can we ensure that the model's prediction meets expectations. ### Using Transformers in Scikit-Learn Using Transformers in Scikit-Learn is quite simple. We can generate a set of simulated data using `make_classification` and then split it into train and test with `train_test_split`. ```Python from sklearn.datasets import make_classification from sklearn.model_selection import train_test_split X, y = make_classification(n_samples=100, n_features=2, n_classes=2, n_redundant=0, n_informative=2, n_clusters_per_class=1, random_state=42) X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42) ``` Let's look at the distribution of the data: ```Python import matplotlib.pyplot as plt plt.scatter(X_train[:, 0], X_train[:, 1], color='red', marker='o') plt.scatter(X_test[:, 0], X_test[:, 1], color='green', marker='s') plt.xlabel('feature_idx_0') plt.ylabel('feature_idx_1') plt.tight_layout() plt.show() ``` ![The distribution of the data before transformation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-7.png) The distribution of the data before scaling. Image by Author Here we're using `StandardScaler` to scale the features. First, initialize the `StandardScaler`, then `fit` it with train data: ```Python from sklearn.preprocessing import StandardScaler scaler = StandardScaler() scaler.fit(X_train) ``` Next, we can `transform` the train data's features with the fitted Transformer: ```Python X_train_std = scaler.transform(X_train) ``` Of course, we could also use `fit_transform` to fit and transform the train data in one go: ```Python X_train_std = scaler.fit_transform(X_train) ``` Then we simply transform the test data without needing to fit it again: ```Python X_test_std = scaler.transform(X_test) ``` After transformation, the distribution of the data remains unchanged, except for the change in scale: ![The distribution of the data after transformation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-8.png) The distribution of the data after scaling. Image by Author Apart from scaling data with tools like [StandardScaler](https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.StandardScaler.html?ref=dataleadsfuture.com#sklearn.preprocessing.StandardScaler) and [MinMaxScaler](https://scikit-learn.org/stable/modules/generated/sklearn.preprocessing.MinMaxScaler.html?ref=dataleadsfuture.com#sklearn.preprocessing.MinMaxScaler), we can also use `PCA`, `SelectKBest`, etc., for dimensionality reduction. For the sake of brevity, I won't delve into these here, but you're welcome to consult the [official documentation](https://scikit-learn.org/stable/modules/feature%5Fselection.html?ref=dataleadsfuture.com#univariate-feature-selection) for more information. --- ## Using Transformers in a Pipeline ### Why use a Pipeline As mentioned earlier, in a machine learning task, we often need to use various Transformers for data scaling and feature dimensionality reduction before training a model. This presents several challenges: - **Code complexity**: For each use of a Transformer, we have to go through initialization, `fit_transform`, and `transform` steps. Missing one step during a transformation could derail the entire training process. - **Data leakage**: As we discussed, for each Transformer, we fit with train data and then transform both train and test data. We must avoid letting the distribution of the test data leak into the train data. - **Code reusability**: A machine learning model includes not only the trained Estimator for prediction but also the data preprocessing steps. Therefore, a machine learning task comprising Transformers and an Estimator should be atomic and indivisible. - **Hyperparameter tuning**: After setting up the steps of machine learning, we need to adjust hyperparameters to find the best combination of Transformer parameter values. Scikit-Learn introduced the `Pipeline` module to solve these issues. ### What is a Pipeline A `Pipeline` is a module in Scikit-Learn that implements the chain of responsibility design pattern. When creating a Pipeline, we use the `steps` parameter to chain together multiple Transformers for initialization: ```Python from sklearn.pipeline import Pipeline from sklearn.decomposition import PCA from sklearn.ensemble import RandomForestClassifier pipeline = Pipeline(steps=[('scaler', StandardScaler()), ('pca', PCA(n_components=2, random_state=42)), ('estimator', RandomForestClassifier(n_estimators=3, max_depth=5))]) ``` The [official documentation](https://scikit-learn.org/stable/modules/compose.html?ref=dataleadsfuture.com#pipeline) points out that the last Transformer must be an Estimator. If you don't need to specify each Transformer's name, you can simplify the creation of a Pipeline with `make_pipeline`: ```Python from sklearn.pipeline import make_pipeline pipeline_2 = make_pipeline(StandardScaler(), PCA(n_components=2, random_state=42), RandomForestClassifier(n_estimators=3, max_depth=5)) ``` ### Understanding the Pipeline's mechanism from the source code We've mentioned the importance of not letting test data variables leak into training data when using each Transformer. This principle is relatively easy to ensure when each data preprocessing step is independent. But what if we integrate these steps using a Pipeline? If we look at the [official documentation](https://scikit-learn.org/stable/modules/compose.html?ref=dataleadsfuture.com#pipeline), we find it simply uses the `fit` method on the entire dataset without explaining how to handle train and test data separately. With this question in mind, I dived into the Pipeline's source code to find the answer. Reading the source code revealed that although Pipeline implements `fit`, `fit_transform`, and `predict` methods, they work differently from regular Transformers. Take the following Pipeline creation process as an example: ```Python from sklearn.pipeline import Pipeline from sklearn.decomposition import PCA from sklearn.ensemble import RandomForestClassifier pipeline = Pipeline(steps=[('scaler', StandardScaler()), ('pca', PCA(n_components=2, random_state=42)), ('estimator', RandomForestClassifier(n_estimators=3, max_depth=5))]) ``` The internal implementation can be represented by the following diagram: ![Internal implementation of the fit and predict methods when called.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/pipeline_transformers-pipeline.webp) Internal implementation of the fit and predict methods when called. Image by Author As you can see, when we call the `fit` method, Pipeline first separates Transformers from the Estimator. For each Transformer, Pipeline checks if there's a `fit_transform` method; if so, it calls it; otherwise, it calls `fit`. For the Estimator, it calls `fit` directly. For the predict method, Pipeline separates Transformers from the Estimator. Pipeline calls each Transformer's `transform` method in sequence, followed by the Estimator's `predict` method. Therefore, when using a Pipeline, we still need to split train and test data. Then we simply call `fit` on the train data and `predict` on the test data. There's a special case when combining Pipeline with `GridSearchCV` for hyperparameter tuning: you don't need to manually split train and test data. I'll explain this in more detail in the best practices section. --- ## Best Practices for Using Transformers and Pipeline in Actual Applications Now that we've discussed the working principles of Transformers and Pipeline, it's time to fulfill the promise made in the title and talk about the best practices when combining Transformers with Pipeline in real projects. ### Combining Pipeline with GridSearchCV for hyperparameter tuning In a machine learning project, selecting the right dataset processing and algorithm is one aspect. After debugging the initial steps, it's time for parameter optimization. Using `GridSearchCV` or `RandomizedSearchCV`, you can try different parameters for the Estimator to find the best fit: ```Python import time from sklearn.model_selection import GridSearchCV pipeline = Pipeline(steps=[('scaler', StandardScaler()), ('pca', PCA()), ('estimator', RandomForestClassifier())]) param_grid = {'pca__n_components': [2, 'mle'], 'estimator__n_estimators': [3, 5, 7], 'estimator__max_depth': [3, 5]} start = time.perf_counter() clf = GridSearchCV(pipeline, param_grid=param_grid, cv=5, n_jobs=4) clf.fit(X, y) # It takes 2.39 seconds to finish the search on my laptop. print(f"It takes {time.perf_counter() - start} seconds to finish the search.") ``` But in machine learning, hyperparameter tuning is not limited to Estimator parameters; it also involves combinations of Transformer parameters. Integrating all steps with Pipeline allows for hyperparameter tuning of every element with different parameter combinations. Note that during hyperparameter tuning, we no longer need to manually split train and test data. `GridSearchCV` will split the data into training and validation sets using [StratifiedKFold](https://scikit-learn.org/stable/modules/cross%5Fvalidation.html?ref=dataleadsfuture.com#stratified-k-fold), which implemented a k-fold cross validation mechanism. ![StratifiedKFold iterative process of splitting train data and test data.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/pipeline_transformers-StratifiedKFold.webp) StratifiedKFold iterative process of splitting train data and test data. Image by Author We can also set the number of folds for cross-validation and choose how many workers to use. The tuning process is illustrated in the following diagram: ![Internal implementation of GridSearchCV hyperparameter tuning.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/pipeline_transformers-GridSearchCV.webp) Internal implementation of GridSearchCV hyperparameter tuning. Image by Author Due to space constraints, I won't go into detail about `GridSearchCV` and `RandomizedSearchCV` here. If you're interested, I can write another article explaining them next time. ### Using the memory parameter to cache Transformer outputs Of course, hyperparameter tuning with `GridSearchCV` can be slow, but that's no worry, Pipeline provides a caching mechanism to speed up the tuning efficiency by caching the results of intermediate steps. When initializing a Pipeline, you can pass in a `memory` parameter, which will cache the results after the first call to `fit` and `transform` for each transformer. If subsequent calls to `fit` and `transform` use the same parameters, which is very likely during hyperparameter tuning, these steps will directly read the results from the cache instead of recalculating, significantly speeding up the efficiency when running the same Transformer repeatedly. The `memory` parameter can accept the following values: - The default is None: caching is not used. - A string: providing a path to store the cached results. - A `joblib.Memory` object: allows for finer-grained control, such as configuring the storage backend for the cache. Next, let's use the previous `GridSearchCV` example, this time adding `memory` to the Pipeline to see how much speed can be improved: ```Python pipeline_m = Pipeline(steps=[('scaler', StandardScaler()), ('pca', PCA()), ('estimator', RandomForestClassifier())], memory='./cache') start = time.perf_counter() clf_m = GridSearchCV(pipeline_m, param_grid=param_grid, cv=5, n_jobs=4) clf_m.fit(X, y) # It takes 0.22 seconds to finish the search with memory parameter. print(f"It takes {time.perf_counter() - start} seconds to finish the search with memory.") ``` As shown, with caching, the tuning process only takes 0.2 seconds, a significant speed increase from the previous 2.4 seconds. ### How to debug Scikit-Learn Pipeline After integrating Transformers into a Pipeline, the entire preprocessing and transformation process becomes a black box. It can be difficult to understand which step the process is currently on. Fortunately, we can solve this problem by adding logging to the Pipeline. We need to create custom transformers to add logging at each step of data transformation. Here's an example of adding logging with Python's standard logging library: First, you need to configure a logger: ```Python import logging from sklearn.base import BaseEstimator, TransformerMixin logging.basicConfig(level=logging.INFO, format='%(asctime)s - %(levelname)s - %(message)s') logger = logging.getLogger() ``` Next, you can create a custom Transformer and add logging within its methods: ```Python class LoggingTransformer(BaseEstimator, TransformerMixin): def __init__(self, transformer): self.transformer = transformer self.real_name = self.transformer.__class__.__name__ def fit(self, X, y=None): logging.info(f"Begin fit: {self.real_name}") self.transformer.fit(X, y) logging.info(f"End fit: {self.real_name}") return self def fit_transform(self, X, y=None): logging.info(f"Begin fit_transform: {self.real_name}") X_fit_transformed = self.transformer.fit_transform(X, y) logging.info(f"End fit_transform: {self.real_name}") return X_fit_transformed def transform(self, X): logging.info(f"Begin transform: {self.real_name}") X_transformed = self.transformer.transform(X) logging.info(f"End transform: {self.real_name}") return X_transformed ``` Then you can use this `LoggingTransformer` when creating your Pipeline: ```Python pipeline_logging = Pipeline(steps=[('scaler', LoggingTransformer(StandardScaler())), ('pca', LoggingTransformer(PCA(n_components=2))), ('estimator', RandomForestClassifier(n_estimators=5, max_depth=3))]) pipeline_logging.fit(X_train, y_train) ``` ![The effect after adding the LoggingTransformer. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/12/image-9.png) The effect after adding the LoggingTransformer. Image by Author When you use `pipeline.fit`, it will call the `fit` and `transform` methods for each step in turn and log the appropriate messages. ### Use passthrough in Scikit-Learn Pipeline In a Pipeline, a step can be set to `'passthrough'`, which means that for this specific step, the input data will pass through unchanged to the next step. This is useful when you want to selectively enable/disable certain steps in a complex pipeline. Taking the code example above, we know that when using `DecisionTree` or `RandomForest`, standardizing the data is unnecessary, so we can use `passthrough` to skip this step. An example would be as follows: ```Python param_grid = {'scaler': ['passthrough'], 'pca__n_components': [2, 'mle'], 'estimator__n_estimators': [3, 5, 7], 'estimator__max_depth': [3, 5]} clf = GridSearchCV(pipeline, param_grid=param_grid, cv=5, n_jobs=4) clf.fit(X, y) ``` ### Reusing the Pipeline After a journey of trials and tribulations, we finally have a well-performing machine learning model. Now, you might consider how to reuse this model, share it with colleagues, or deploy it in a production environment. However, the result of a model's training includes not only the model itself but also the various data processing steps, which all need to be saved. Using `joblib` and Pipeline, we can save the entire training process for later use. The following code provides a simple example: ```Python from joblib import dump, load # save pipeline dump(pipeline, 'model_pipeline.joblib') # load pipeline loaded_pipeline = load('model_pipeline.joblib') # predict with loaded pipeline loaded_predictions = loaded_pipeline.predict(X_test) ``` --- ## Conclusion The [official Scikit-Learn documentation](https://scikit-learn.org/stable/user%5Fguide.html?ref=dataleadsfuture.com) is among the best I've seen. By learning its contents, you can master the basics of machine learning applications. However, when using Scikit-Learn in real projects, we often encounter various details that the official documentation may not cover. How to correctly combine Transformers with Pipeline is one such case. In this article, I introduced why to use Transformers and some typical application scenarios. Then, I interpreted the working principle of Pipeline from the source code level and completed the reasonable use case when applied to train and test datasets. Finally, for each stage of a real machine learning project, I introduced the best practices of combining Transformers with Pipeline based on my work experience. I hope this article can help you. If you have any questions, please leave me a message, and I will try my best to answer them. ### How to Optimize Multidimensional Numpy Array Operations with Numexpr URL: https://www.dataleadsfuture.com/how-to-optimize-multidimensional-numpy-array-operations-with-numexpr/ Last updated: 2025-03-26T02:14:20.000Z This is a relatively brief article. In it, I will use a real-world scenario as an example to explain how to use [Numexpr expressions](https://numexpr.readthedocs.io/en/latest/user%5Fguide.html?ref=dataleadsfuture.com#supported-functions) in multidimensional Numpy arrays to achieve substantial performance improvements. There aren't many articles explaining how to use Numexpr in multidimensional Numpy arrays and how to use Numexpr expressions, so I hope this one will help you. --- ## Introduction Recently, while reviewing some of my old work, I stumbled upon this piece of code: ```Python def predict(X, w, b): z = np.dot(X, w) y_hat = sigmoid(z) y_pred = np.zeros((y_hat.shape[0], 1)) for i in range(y_hat.shape[0]): if y_hat[i, 0] < 0.5: y_pred[i, 0] = 0 else: y_pred[i, 0] = 1 return y_pred ``` This code transforms prediction results from probabilities to classification results of 0 or 1 in the logistic regression model of machine learning. But heavens, who would use a `for loop` to iterate over Numpy ndarray? You can foresee that when the data reaches a certain amount, it will not only occupy a lot of memory, but the performance will also be inferior. That's right, the person who wrote this code was me when I was younger. With a sense of responsibility, I plan to rewrite this code with the Numexpr library today. Along the way, I will show you how to use Numexpr and Numexpr's `where` expression in multidimensional Numpy arrays to achieve significant performance improvements. --- ## Code Implementation If you are not familiar with the basic usage of Numexpr, you can refer to this article: [Exploring Numexpr: A Powerful Engine Behind PandasEnhancing your data analysis performance with Python’s Numexpr and Pandas’ eval/query functions![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/format/png/2024/01/color.webp)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w1200/2024/01/Numexpr_Vladivostok.webp)](https://www.dataleadsfuture.com/exploring-numexpr-a-powerful-engine-behind-pandas/) This article uses a real-world example to demonstrate the specific usage of Numexpr's API and expressions in Numpy and Pandas. > `where(bool, number1, number2): number` \- number1 if the bool condition is true, number2 otherwise. The above is the usage of the where expression in Numpy. When dealing with matrix data, you may used to using Pandas `DataFrame`. But since the `eval` method of Pandas does not support the `where` expression, you can only choose to use Numexpr in multidimensional Numpy ndarray. Don't worry, I'll explain it to you right away. Before starting, we need to import the necessary packages and implement a `generate_ndarray` method to generate a specific size ndarray for testing: ```Python from typing import Callable import time import numpy as np import numexpr as ne import matplotlib.pyplot as plt rng = np.random.default_rng(seed=4000) def generate_ndarray(rows: int) -> np.ndarray: result_array = rng.random((rows, 1)) return result_array ``` First, we generate a matrix of 200 rows to see if it is the test data we want: ```Python In: arr = generate_ndarray(200) print(f"The dimension of this array: {arr.ndim}") print(f"The shape of this array: {arr.shape}") Out: The dimension of this array: 2 The shape of this array: (200, 1) ``` To be close to the actual situation of the logistic regression model, we generate an ndarray of the shape `(200, 1)`an. Of course, you can also test other shapes of ndarray according to your needs. Then, we start writing the specific use of Numexpr in the `numexpr_to_binary method`: - First, we use the index to separate the columns that need to be processed. - Then, use the where expression of Numexpr to process the values. - Finally, merge the processed columns with other columns to generate the required results. Since the ndarray's shape here is `(200, 1)`, there is only one column, so I add a new dimension. The code is as follows: ```Python def numexpr_to_binary(np_array: np.ndarray) -> np.ndarray: temp = np_array[:, 0] temp = ne.evaluate("where(temp<0.5, 0, 1)") return temp[:, np.newaxis] ``` We can test the result with an array of 10 rows to see if it is what I want: ```Python arr = generate_ndarray(10) result = numexpr_to_binary(arr) mapping = np.column_stack((arr, result)) mapping ``` ![I test an array of 10 rows and the result is what I want. ](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/10/image.png) I test an array of 10 rows and the result is what I want. Image by Author Look, the match is correct. Our task is completed. The entire process can be demonstrated with the following figure: ![The entire process of how Numexpr transforms the multidimensional ndarray.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/10/numpy_arrays-numexpr.drawio.png) The entire process of how Numexpr transforms the multidimensional ndarray. Image by Author --- ## Performance Comparison After the code implementation, we need to compare the Numexpr implementation version with the previous `for each` implementation version to confirm that there has been a performance improvement. First, we implement a `numexpr_example` method. This method is based on the implementation of Numexpr: ```Python def numexpr_example(rows: int) -> np.ndarray: orig_arr = generate_ndarray(rows) the_result = numexpr_to_binary(orig_arr) return the_result ``` Then, we need to supplement a `for_loop_example` method. This method refers to the original code I need to rewrite and is used as a performance benchmark: ```Python def for_loop_example(rows: int) -> np.ndarray: the_arr = generate_ndarray(rows) for i in range(the_arr.shape[0]): if the_arr[i][0] < 0.5: the_arr[i][0] = 0 else: the_arr[i][0] = 1 return the_arr ``` Then, I wrote a test method `time_method`. This method will generate data from 10 to 10 to the 9th power rows separately, call the corresponding method, and finally save the time required for different data amounts: ```Python def time_method(method: Callable): time_dict = dict() for i in range(9): begin = time.perf_counter() rows = 10 ** i method(rows) end = time.perf_counter() time_dict[i] = end - begin return time_dict ``` We test the numexpr version and the `for_loop` version separately, and use `matplotlib` to draw the time required for different amounts of data: ```Python t_m = time_method(for_loop_example) t_m_2 = time_method(numexpr_example) plt.plot(t_m.keys(), t_m.values(), c="red", linestyle="solid") plt.plot(t_m_2.keys(), t_m_2.values(), c="green", linestyle="dashed") plt.legend(["for loop", "numexpr"]) plt.xlabel("exponent") plt.ylabel("time") plt.show() ``` ![The Numexpr version of the implementation has a huge performance improvement.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/10/image-1.png) The Numexpr version of the implementation has a huge performance improvement. Image by Author It can be seen that when the number of rows of data is greater than 10 to the 6th power, the Numexpr version of the implementation has a huge performance improvement. --- ## Conclusion After explaining the basic usage of Numexpr in the previous article, this article uses a specific example in actual work to explain how to use Numexpr to rewrite existing code to obtain performance improvement. This article mainly uses two features of Numexpr: 1. Numexpr allows calculations to be performed in a vectorized manner. 2. During the calculation of Numexpr, no new arrays will be generated, thereby significantly reducing memory usage. Thank you for reading. If you have other solutions, please feel free to leave a message and discuss them with me. ### Exploring Numexpr: A Powerful Engine Behind Pandas URL: https://www.dataleadsfuture.com/exploring-numexpr-a-powerful-engine-behind-pandas/ Last updated: 2025-03-26T02:13:49.000Z This article will introduce you to the Python library [Numexpr](https://numexpr.readthedocs.io/en/latest/intro.html?ref=dataleadsfuture.com), a tool that boosts the computational performance of `Numpy Arrays`. The `eval` and `query` methods of Pandas are also based on this library. This article also includes a hands-on weather data analysis project. By reading this article, you will understand the principles of Numexpr and how to use this powerful tool to speed up your calculations in reality. --- ## Introduction ### Recalling Numpy Arrays In a previous article discussing `Numpy Arrays`, I used a library example to explain why Numpy's Cache Locality is so efficient: [Python Lists Vs. NumPy Arrays: A Deep Dive into Memory Layout and Performance BenefitsExploring allocation differences and efficiency gains![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_30niIoNGPitrWYdT.jpg)](https://www.dataleadsfuture.com/python-lists-vs-numpy-arrays-a-deep-dive-into-memory-layout-and-performance-benefits/) Each time you go to the library to search for materials, you take out a few books related to the content and place them next to your desk. This way, you can quickly check related materials without having to run to the shelf each time you need to read a book. This method saves a lot of time, especially when you need to consult many related books. In this scenario, the shelf is like your memory, the desk is equivalent to the CPU's L1 cache, and you, the reader, are the CPU's core. ![When the CPU accesses RAM, the cache loads the entire cache line into the high-speed cache.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-74.png) When the CPU accesses RAM, the cache loads the entire cache line into the high-speed cache. Image by Author ### The limitations of Numpy Suppose you are unfortunate enough to encounter a demanding professor who wants you to take out Shakespeare and Tolstoy's works for a cross-comparison. At this point, taking out related books in advance will not work well. First, your desk space is limited and cannot hold all the books of these two masters at the same time, not to mention the reading notes that will be generated during the comparison process. Second, you're just one person, and comparing so many works would take too long. It would be nice if you could find a few more people to help. This is the current situation when we use Numpy to deal with large amounts of data: - The number of elements in the Array is too large to fit into the CPU's L1 cache. - Numpy's element-level operations are single-threaded and cannot utilize the computing power of multi-core CPUs. What should we do? Don't worry. When you really encounter a problem with too much data, you can call on our protagonist today, `Numexpr`, to help. --- ## Understanding Numexpr: What and Why ### How it works When Numpy encounters large arrays, element-wise calculations will experience two extremes. Let me give you an example to illustrate. Suppose there are two large Numpy ndarrays: ```Python import numpy as np import numexpr as ne a = np.random.rand(100_000_000) b = np.random.rand(100_000_000) ``` When calculating the result of the expression `a**5 + 2 * b`, there are generally two methods: One way is Numpy's vectorized calculation method, which uses two temporary arrays to store the results of `a**5` and `2*b` separately. ```Python In: %timeit a**5 + 2 * b Out:2.11 s ± 31.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) ``` At this time, you have four arrays in your memory: `a`, `b`, `a**5`, and `2 * b`. This method will cause a lot of memory waste. Moreover, since each Array's size exceeds the CPU cache's capacity, it cannot use it well. Another way is to traverse each element in two arrays and calculate them separately. ```Python c = np.empty(100_000_000, dtype=np.uint32) def calcu_elements(a, b, c): for i in range(0, len(a), 1): c[i] = a[i] ** 5 + 2 * b[i] %timeit calcu_elements(a, b, c) Out: 24.6 s ± 48.2 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) ``` This method performs even worse. The calculation will be very slow because it cannot use vectorized calculations and only partially utilize the CPU cache. ### Numexpr's calculation Numexpr commonly uses only one `evaluate` method. This method will receive an expression string each time and then compile it into bytecode using Python's `compile` method. Numexpr also has a virtual machine program. The virtual machine contains multiple vector registers, each using a chunk size of 4096. When Numexpr starts to calculate, it sends the data in one or more registers to the CPU's L1 cache each time. This way, there won't be a situation where the memory is too slow, and the CPU waits for data. At the same time, Numexpr's virtual machine is written in C, removing Python's GIL. It can utilize the computing power of multi-core CPUs. So, Numexpr is faster when calculating large arrays than using Numpy alone. We can make a comparison: ```Python In: %timeit ne.evaluate('a**5 + 2 * b') Out: 258 ms ± 14.4 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) ``` ### Summary of Numexpr's working principle Let's summarize the working principle of Numexpr and see why Numexpr is so fast: **Executing bytecode through a virtual machine.** Numexpr uses bytecode to execute expressions, which can fully utilize the [branch prediction](https://en.wikipedia.org/wiki/Branch%5Fpredictor?ref=dataleadsfuture.com) ability of the CPU, which is faster than using Python expressions. **Vectorized calculation.** Numexpr will use [SIMD (Single Instruction, Multiple Data)](https://en.wikipedia.org/wiki/Single%5Finstruction,%5Fmultiple%5Fdata?ref=dataleadsfuture.com) technology to improve computing efficiency significantly for the same operation on the data in each register. **Multi-core parallel computing.** Numexpr's virtual machine can decompose each task into multiple subtasks. They are executed in parallel on multiple CPU cores. **Less memory usage.** Unlike Numpy, which needs to generate intermediate arrays, Numexpr only loads a small amount of data when necessary, significantly reducing memory usage. ![Workflow diagram of Numexpr.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/numexpr.drawio.png) Workflow diagram of Numexpr. Image by Author --- ## Numexpr and Pandas: A Powerful Combination You might be wondering: We usually do data analysis with pandas. I understand the performance improvements Numexpr offers for Numpy, but does it have the same improvement for Pandas? The answer is Yes. The `eval` and `query` methods in pandas are implemented based on Numexpr. Let's look at some examples: ### Pandas.eval for Cross-DataFrame operations When you have multiple pandas DataFrames, you can use `pandas.eval` to perform operations between DataFrame objects, for example: ```Python import pandas as pd nrows, ncols = 1_000_000, 100 df1, df2, df3, df4 = (pd.DataFrame(rng.random((nrows, ncols))) for i in range(4)) ``` If you calculate the sum of these DataFrames using the traditional pandas method, the time consumed is: ```Python In: %timeit df1+df2+df3+df4 Out: 1.18 s ± 65.1 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) ``` You can also use `pandas.eval` for calculation. The time consumed is: ```Python In: %timeit pd.eval('df1 + df2 + df3 + df4') Out: 452 ms ± 29.4 ms per loop (mean ± std. dev. of 7 runs, 1 loop each) ``` The calculation of the eval version can improve performance by 50%, and the results are precisely the same: ```Python In: np.allclose(df1+df2+df3+df4, pd.eval('df1+df2+df3+df4')) Out: True ``` ### DataFrame.eval for column-level operations Just like `pandas.eval`, DataFrame also has its own `eval` method. We can use this method for column-level operations within DataFrame, for example: ```Python df = pd.DataFrame(rng.random((1000, 3)), columns=['A', 'B', 'C']) result1 = (df['A'] + df['B']) / (df['C'] - 1) result2 = df.eval('(A + B) / (C - 1)') ``` The results of using the traditional pandas method and the `eval` method are precisely the same: ```Python In: np.allclose(result1, result2) Out: True ``` Of course, you can also directly use the `eval` expression to add new columns to the DataFrame, which is very convenient: ```Python df.eval('D = (A + B) / C', inplace=True) df.head() ``` ![Directly use the eval expression to add new columns.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/image-1.png) Directly use the eval expression to add new columns. Image by Author ### Using DataFrame.query to quickly find data If the `eval` method of DataFrame executes comparison expressions, the returned result is a boolean result that meets the conditions. You need to use `Mask Indexing` to get the desired data: ```Python mask = df.eval('(A < 0.5) & (B < 0.5)') result1 = df[mask] result1 ``` ![When filtering data only with DataFrame.query, it is necessary to use a boolean mask.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/image-2.png) When filtering data only with DataFrame.query, it is necessary to use a boolean mask. Image by Author The `DataFrame.query` method encapsulates this process, and you can directly obtain the desired data with the `query` method: ```Python In: result2 = df.query('A < 0.5 and B < 0.5') np.allclose(result1, result2) Out: True ``` When you need to use scalars in expressions, you can use the `@` symbol to indicate: ```Python In: Cmean = df['C'].mean() result1 = df[(df.A < Cmean) & (df.B < Cmean)] result2 = df.query('A < @Cmean and B < @Cmean') np.allclose(result1, result2) Out: True ``` --- ## Practical Example: Using Numexpr and Pandas in Real-World Scenarios In all articles explaining Numexpr, examples are made using synthetic data. This feeling is not good and may cause you to not know how to use this powerful library to complete tasks after reading the article. Therefore, in this article, I will take a weather data analysis project as an example to explain how we should use Numexpr to process large datasets in actual work. ### Project Goal After a hot summer, I really want to see if there is such a place where the climate is pleasant in summer and suitable for me to escape the heat. This place should meet the following conditions: 1. In the summer: 2. The daily average temperature is between 18 degrees Celsius and 22 degrees Celsius; 3. The diurnal temperature difference is between 4 degrees Celsius and 6 degrees Celsius; 4. The average wind speed in kmh is between 6 and 10\. It would feel nice to have a breeze blowing on me. ### Data preparation This time, I used the global major city weather data provided by the [Meteostat JSON API](https://dev.meteostat.net/api/stations/daily.html?ref=dataleadsfuture.com#endpoint). The data is licensed under the [Creative Commons Attribution-NonCommercial 4.0 International Public License (CC BY-NC 4.0)](https://dev.meteostat.net/terms.html?ref=dataleadsfuture.com#availability) and can be used commercially. I used the parquet dataset integrated on [Kaggle](https://www.kaggle.com/datasets/guillemservera/global-daily-climate-data?select=daily%5Fweather.parquet&ref=dataleadsfuture.com) based on the Meteostat JSON API for convenience. I used version 2.0 of pandas. The `pandas.read_parquet` method of this version can easily read parquet data. But before reading, you need to install `Pyarrow` and `Fastparquet`. ```Bash conda install pyarrow ``` ```Bash conda install fastparquet ``` ### Data analysis After the preliminary preparations, we officially entered the data analysis process. First, I read the data into memory and then look at the situation of this dataset: ```Python import os from pathlib import Path import pandas as pd root = Path(os.path.abspath("")).parents[0] data = root/"data" df = pd.read_parquet(data/"daily_weather.parquet") df.info() ``` ![Overview of the dataset's metadata.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/image-3.png) Overview of the dataset's metadata. Image by Author As shown in the figure, this dataset contains 13 fields. According to the goal of this project, I plan to use the fields of `city_name`, `season`, `min_temp_c`, `max_temp_c`, `avg_wind_speed_kmh`. Next, I first remove the data in the corresponding fields that contain empty values for subsequent calculations, and then select the desired fields to form a new DataFrame: ```Python sea_level_not_null = df.dropna(subset=['min_temp_c', 'max_temp_c', 'avg_wind_speed_kmh'] , how='any') sample = sea_level_not_null[['city_name', 'season', 'min_temp_c', 'max_temp_c', 'avg_wind_speed_kmh']] ``` Since I need to calculate the average temperature and temperature difference, I use the `Pandas.eval` method to directly calculate the new indicators on the DataFrame: ```Python sample.eval('avg_temp_c = (max_temp_c + min_temp_c) / 2', inplace=True) sample.eval('diff_in_temp = max_temp_c - min_temp_c', inplace=True) ``` Then, average a few indicators by `city_name` and `season`: ```Python sample = sample.groupby(['city_name', 'season'])\ [['min_temp_c', 'max_temp_c', 'avg_temp_c', 'diff_in_temp', 'avg_wind_speed_kmh']]\ .mean().round(1).reset_index() sample ``` ![Results after data cleaning and metric calculation.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/image-4.png) Results after data cleaning and metric calculation. Image by Author Finally, according to the goal of the project, I use `DataFrame.query` to filter the dataset: ```Python sample.query('season=="Summer" \ & 18 < avg_temp_c < 22 \ & 4 < diff_in_temp < 6 \ & 6 < avg_wind_speed_kmh < 10') ``` ![Finally, we obtained the only result that met the criteria.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/09/image-5.png) Finally, we obtained the only result that met the criteria. Image by Author The final result is out. Only one city meets my requirements: Vladivostok, a non-freezing port in the east of Russia. It is indeed an excellent place to escape the heat! --- ## Best Practices and Takeaways After explaining the project practice of Numexpr, as usual, I will explain some of the best practices of Numexpr combined with my own work experience for you. ### Avoid overuse Although Numexpr and pandas `eval` have significant performance advantages when handling large data sets. However, dealing with small data sets is not faster than regular operations. Therefore, you should choose whether to use Numexpr based on the size and complexity of the data. And my experience is to use it when you feel the need, as small datasets won't slow things down too much anyway. ### The use of the eval function is limited The `eval` function does not support all Python and pandas operations. Therefore, before using it, you should consult the [documentation](https://pandas.pydata.org/docs/user%5Fguide/enhancingperf.html?ref=dataleadsfuture.com#expression-evaluation-via-eval) to understand what operations eval supports. ### Be careful when handling strings Although I used `season="Summer"` to filter the dataset in the project practice, the `eval` function is not very fast when dealing with strings. If you have a lot of string operations in your project, you need to consider other ways. ### Be mindful of memory usage Although Numexpr no longer generates intermediate arrays, large datasets will occupy a lot of memory. For example, the dataset occupies 2.6G of memory in my project example. At this time, you have to be very careful to avoid the program crashing due to insufficient memory. ### Use the appropriate data type This point is detailed in the [official documentation](https://numexpr.readthedocs.io/en/latest/user%5Fguide.html?ref=dataleadsfuture.com#datatypes-supported-internally), so I won't repeat it here. ### Use the inplace parameter when needed Using the `inplace` parameter of the `DataFrame.eval` method can directly modify the original dataset, avoiding generating a new dataset and occupying a lot of memory. Of course, doing so will lead to modifications to the original dataset, so please be careful. --- ## Conclusion In this article, I brought a comprehensive tutorial on Numexpr, including: The applicable scenarios of Numexpr, the effect of performance improvement, and its working principle. The `eval` and `query` methods in Pandas are also based on Numexpr. It will bring great convenience and performance improvement to your pandas' operations if used appropriately. Through a global weather data analysis project, I demonstrated how to use pandas' `eval` and `query` methods in practice. As always, combined with my work experience, I introduced the best practices of Numexpar and the eval method of pandas. Thank you for reading. If you have any questions, please leave a message in the comment area, and I will answer in time. ### Seaborn 0.12: An Insightful Guide to the Objects Interface and Declarative Graphics URL: https://www.dataleadsfuture.com/seaborn-0-12-an-insightful-guide-to-the-objects-interface-and-declarative-graphics/ Last updated: 2025-12-19T07:18:04.000Z This article aims to introduce the objects interface feature in [Seaborn 0.12](https://seaborn.pydata.org/whatsnew/index.html?ref=dataleadsfuture.com#v0-12), including the concept of declarative graphic syntax, and a practical visualization project to showcase the usage of the objects interface. By the end of this article, you'll have a clear understanding of the advantages and limitations of [Seaborn's objects interface API](https://seaborn.pydata.org/tutorial/objects%5Finterface.html?ref=dataleadsfuture.com). And you will be able to use Seaborn for data analysis projects more easily. --- ## Introduction Imagine you're creating a data visualization chart using Python. You have to instruct the computer every step of the way: select a dataset, create a figure, set the color, add labels, adjust the size, etc... Then you realize your code is getting longer and more complex, and all you wanted was to quickly visualize your data. It's like going to the grocery store and having to specify every item's location, color, size, and shape, instead of just telling the shop assistant what you need. Not only is this time-consuming, but it can also feel tiring. However, Seaborn 0.12's new feature—the objects interface—and its use of declarative graphic syntax is like having a shop assistant who understands you. You just need to tell it what you need to do, and it will find everything for you. You no longer need to instruct it every step of the way. You just need to tell it what kind of result you want. In this article, I'll guide you through using the objects interface, this new feature that makes your data visualization process more effortless, flexible, and enjoyable. Let's get started! --- ## Seaborn API: Then and Now Before diving into the objects interface API, let's systematically look at the differences between the Seaborn API of earlier versions and the 0.12 version. ### The original API Many readers might have been intimidated by Matplotlib's complex API documentation when learning Python data visualization. Seaborn simplifies this by wrapping and streamlining Matplotlib's API, making the learning curve gentler. Seaborn doesn't just offer high-level encapsulation of Matplotlib; it also categorizes all charts into relational, distributional, and categorical scenarios. ![Overview of Seaborn's original API design.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/seaborn-past-api.drawio.png) Overview of Seaborn's original API design. Image by Author You should comprehensively understand Seaborn's API through this diagram and know when to use which chart. For example, a `histplot` representing data distribution would fall under the distribution chart category. In contrast, a `violinplot` representing data features by category would be classified as a categorical chart. Aside from vertical categorization, Seaborn also performs horizontal categorization: `Figure-level` and `axes-level`. According to the [official website](https://seaborn.pydata.org/tutorial/function%5Foverview.html?ref=dataleadsfuture.com#figure-level-vs-axes-level-functions), `axes-level` charts are drawn on `matplotlib.pyplot.axes` and can only draw one figure. In contrast, `Figure-level` charts use Matplotlib's `FacetGrid` to draw multiple charts in one figure, facilitating easy comparison of similar data dimensions. However, even though Seaborn's API significantly simplifies chart drawing through encapsulating Matplotlib, creating an individual-specific chart still requires complex configurations. For example, if I use Seaborn's built-in `penguins` dataset to draw a `histplot`, the code is as follows: ```Python sns.histplot(penguins, x="flipper_length_mm", hue="species"); ``` ![The original way of drawing a histplot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/original_hist_chart.png) The original way of drawing a histplot. Image by Author And when I use the same dataset to draw a `kdeplot`, the code is as follows: ```Python sns.kdeplot(penguins, x="flipper_length_mm", fill=True, hue="species"); ``` ![The original way of drawing a kdeplot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/original_kde_chart.png) The original way of drawing a kdeplot. Image by Author Except for the chart API, the rest of the configurations are identical. This is like telling the chef I want to use lamb chops and onions to make a lamb soup and specifying the cooking steps. When I want to use these ingredients to make a roasted lamb chop, I have to tell the chef about the ingredients and the cooking steps all over again. Not only is it inefficient, but it also needs more flexibility. That's why Seaborn introduced the objects interface API in its 0.12 version. This declarative graphic syntax dramatically improves the process of creating a chart. ### The objects Interface API Before we start with the objects interface API, let's take a high-level look at it to better understand the drawing process. Unlike the original Seaborn API, which organizes the drawing API by classification, the objects interface API collects the API by a drawing pipeline. The objects interface API divides the drawing into multiple stages, such as data binding, layout, presentation, customization, etc. ![Overview of Seaborn's objects interface API design.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/seaborn-objects-interface.drawio--1-.png) Overview of Seaborn's objects interface API design. Image by Author The data binding and presentation stages are necessary, while other stages are optional. Also, since the stages are independent, each stage can be reused. Following the previous example of the hist and kde plots: To use the objects interface to draw, we first need to bind the data: ```Python p = so.Plot(penguins, x="flipper_length_mm", color="species") ``` From this line of code, we can see that the objects interface uses the `so.Plot` class for data binding. Also, compared to the original API that uses the incomprehensible `hue` parameter, it uses the `color` parameter to bind the `species` dimension directly to the chart color, making the configuration more intuitive. Finally, this line of code returns a `p` instance that can be reused to draw a chart. Next, let's draw a `histplot`: ```Python p.add(so.Bars(), so.Hist()) ``` ![Use objects interface API to draw a histplot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_hist_chart.png) Use objects interface API to draw a histplot. Image by Author This line of code shows that the drawing stage does not need to rebind the data. We just need to tell the `add` method what to draw: `so.Bars()`, and how to calculate it: `so.Hist()`. The `add` method also returns a copy of the `Plot` instance, so any adjustments in the `add` method will not affect the original data binding. The `p` instance can still be reused. Therefore, we continue to call the `p.add()` method to draw a `kdeplot`: ```Python p.add(so.Area(), so.KDE()) ``` ![Use objects interface API to draw a kdeplot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_kde_chart.png) Use objects interface API to draw a kdeplot. Image by Author Since `KDE` is a way of statistic, `so.KDE()` is called on the `stat` parameter here. And since the `kdeplot` itself is an area plot, `so.Area()` is used for drawing. We reused the `p` instance bound to the data, so there is no need to tell the chef how to cook each dish, but to directly say what we want. Isn't it much more concise and flexible? --- ## Unpacking the Objects Interface with Examples Next, see how some common charts are written using the original Seaborn API and the objects interface API. Before we start, we need to import the necessary libraries: ```Python %matplotlib inline import matplotlib.pyplot as plt import seaborn as sns import seaborn.objects as so import pandas as pd sns.set() penguins = sns.load_dataset('penguins') ``` ### Bar chart In the original API, to draw a bar chart, the code is as follows: ```Python sns.barplot(penguins, x="island", y="body_mass_g", hue="species"); ``` ![The original way of drawing a bar chart.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/original_bar_chart.png) The original way of drawing a bar chart. Image by Author In the objects interface, to draw a bar chart, the code is as follows: ```Python ( so.Plot(penguins, x="island", y="body_mass_g", color="species") .add(so.Bar(), so.Dodge()) ) ``` ![Use objects interface to draw a bar chart.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_bar_chart.png) Use objects interface to draw a bar chart. Image by Author ### Scatter plot In the original API, to draw a scatter plot, the code is as follows: ```Python sns.relplot(penguins, x="bill_length_mm", y="bill_depth_mm", hue="species"); ``` ![In the original way, we use relplot to draw a scatter plot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/original_scatter_plot.png) In the original way, we use relplot to draw a scatter plot. Image by Author In the objects interface, to draw a scatter plot, the code is as follows: ```Python ( so.Plot(penguins, x="bill_length_mm", y="bill_depth_mm", color="species") .add(so.Dots()) ) ``` ![When using objects interface, we use so.Dots to draw a scatter plot.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_scatter_plot.png) When using objects interface, we use so.Dots to draw a scatter plot. Image by Author You may think that after comparing the drawing of the two APIs, it doesn't seem like the objects interface is too special either. Don't worry. Let's take a look at the advanced usage of the objects interface. ### Advanced usage Suppose we use Seaborn's `tips` dataset. ```Python tips = sns.load_dataset("tips") ``` I want to use a bar chart to see the average tip for different dates and mark the values on the chart. The chart I want is shown below: ![A bar chart with text to show the values.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_bar_chart_text.png) A bar chart with text to show the values. Image by Author Before we start drawing, we need to process the `tips` dataset to calculate the average value for each day. ```Python day_mean = tips[['day', 'tip']].groupby('day').mean().round(2).reset_index() ``` Then, we can use the objects interface to draw: ```Python ( day_mean .pipe(so.Plot, y="day", x="tip", text="tip") .add(so.Bar(width=.5)) .add(so.Text(color='w', halign="right")) ) ``` We use two tricks here: First, we call the `pipe` method on the `dataframe` to enable chain code calls. Second, we can reuse the instance of `so.Plot`, and only bind the data once to draw multiple graphs. Then, let's see how the code would be written using the original API: ```Python ax = sns.barplot(day_mean, x="tip", y="day") for p in ax.patches: width = p.get_width() ax.text(width, p.get_y() + p.get_height()/2, '{:1.2f}'.format(width), ha="right", va="center") plt.show() ``` As you can see, the original code is much more complex: First, draw a horizontal bar chart. Then use iteration to draw the corresponding values on each bar. In comparison, doesn't the objects interface seem simpler and more flexible? --- ## Applying the Objects Interface to Real-World Data Next, to help everyone deepen their memory and master the usage of the objects interface systematically, I plan to lead everyone to practice in an actual data visualization project. In this project, I plan to visually explore the data of New York City's shared bicycle system to understand the usage of the city's shared bicycles and help enterprises operate better. ### Data source We will use the Citi Bike Sharing dataset from Citibikenyc in this project. You can find the dataset here: [https://citibikenyc.com/system-data](https://citibikenyc.com/system-data?ref=dataleadsfuture.com) To facilitate the following coding process, I cleaned and merged the data in this dataset and finally synthesized one data set. ### Data preprocessing Before we begin, we should understand the fields included in this dataset, which can be achieved by executing the following code: ```Python citibike = pd.read_csv("../data/CitiBike-2021-combined.csv", index_col="ID") citibike.info() ``` ```Bash Data columns (total 15 columns): # Column Non-Null Count Dtype --- ------ -------------- ----- 0 Trip Duration 735502 non-null int64 1 Start Time 735502 non-null datetime64[ns] 2 Stop Time 735502 non-null datetime64[ns] 3 Start Station ID 735502 non-null int64 4 Start Station Name 735502 non-null object 5 Start Station Latitude 735502 non-null float64 6 Start Station Longitude 735502 non-null float64 7 End Station ID 735502 non-null int64 8 End Station Name 735502 non-null object 9 End Station Latitude 735502 non-null float64 10 End Station Longitude 735502 non-null float64 11 Bike ID 735502 non-null int64 12 User Type 735502 non-null object 13 Birth Year 735502 non-null int64 14 Gender 735502 non-null object dtypes: datetime64[ns](2), float64(4), int64(8), object(6) memory usage: 117.8+ MB ``` This dataset contains 15 fields, and since our goal is to understand the usage of shared bicycles in the city, all 15 fields will be helpful for us. Also, to facilitate the analysis of the use of shared bicycles in different months of each year, as well as on weekdays and non-working days of each week, I need to generate two fields for the dataset: `Start Month` and `Day Of Week`: ```Python citibike['Start Time'] = pd.to_datetime(citibike['Start Time']) citibike['Stop Time'] = pd.to_datetime(citibike['Stop Time']) citibike['Day Of Week'] = citibike['Start Time'].dt.day_of_week citibike['Start Month'] = citibike['Start Time'].dt.month day_dict = {0: 'Mon', 1: 'Tue', 2: 'Wen', 3: 'Thu', 4: 'Fri', 5: 'Sat', 6: 'Sun'} citibike['Day Of Week'] = citibike['Day Of Week'].replace(day_dict) ``` To facilitate display, I will convert the `Gender` field into text gender, convert the `Birth Year` into `Decade`, and change `Trip Duration` from seconds to minutes: ```Python citibike['Gender'] = citibike['Gender'].replace({0: 'Unknown', 1: 'Male', 2: 'Female'}) citibike['Decade'] = (citibike['Birth Year'] // 10 * 10).astype(str) + 's' citibike['Duration_Min'] = citibike['Trip Duration'] // 60 ``` Finally, since the original dataset is large, we only need to find out the distribution of the data, so I will sample the dataset for easier and faster drawing: ```Python citibike_sample = citibike.sample(n=10000, random_state=1701) ``` ### Visual analysis **Remember, the purpose of data visualization is not just to display data, but to excavate the story behind the data.** In this project, I expect to understand under what circumstances users will use shared bicycles, to facilitate the distribution of bicycles or carry out corresponding promotions. First, I want to see in which season people are more inclined to use shared bicycles. Since I want to see the total amount of data by month, I directly use the original dataset for drawing. But to speed up the drawing, I aggregate the data in the `dataframe` and then call the pipeline using the `pipe` method. ```Python ( citibike.groupby('Start Month').size().reset_index(name="Count") .pipe(so.Plot, x="Start Month", y="Count") .add(so.Line(marker='o', edgecolor='w')) .add(so.Text(valign='bottom'), text='Count') ) ``` ![View shared bike usage by month.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_distribution_month.png) View shared bike usage by month. Image by Author The chart shows that bicycles have more uses in March and October of each year. This indicates that people are more willing to ride bikes in a mild climate. Next, I want to see which days of the week people use shared bicycles more. Since we only need to see a proportion here, I use the sampled dataset and set a `proportion` in `so.Hist()`. ```Python ( so.Plot(citibike_sample, x="Day Of Week", color="Gender") .scale(x=so.Nominal(order=['Mon', 'Tue', 'Wen', 'Thu', 'Fri', 'Sat', 'Sun'])) .add(so.Bar(), so.Hist(stat="proportion"), so.Dodge()) ) ``` ![Which days of the week do people use shared bicycles more.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_weekday_usage.png) Which days of the week do people use shared bicycles more? Image by Author Both males and females use shared bicycles more on weekdays, probably for commuting to work. But we also found that users with 'Unknown' gender use shared bicycles more on weekends. Why is this the case? We can continue to explore. Next, I want to see the proportion of cycling duration in different gender situations. Here I will draw a histogram for each gender separately and use `facet` for layout. To eliminate the interference generated by anomalous data, I only took data within one standard deviation of the average riding time for reference. ```Python mean = citibike_sample["Duration_Min"].mean() std = citibike_sample["Duration_Min"].std() citibike_filterd = citibike_sample.query("(Duration_Min > @mean - @std) and (Duration_Min < @mean + @std)") ( so.Plot(citibike_filterd, x="Duration_Min") .facet(col="Gender") .layout(size=(6,3)) .add(so.Bars(), so.Hist(stat="proportion")) ) ``` ![A histogram for each gender separately to show the proportion of cycling duration.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_duration_by_gender.png) A histogram for each gender separately to show the proportion of cycling duration. Image by Author The chart shows that the cycling duration of males and females conforms to our cognition. Still, the cycling duration of users with the 'Unknown' gender seems more evenly distributed, indicating that cycling is more casual and lacks purpose. Fourth, I want to understand the proportion of cycling duration by membership category: ```Python ( so.Plot(citibike_filterd, x="Duration_Min") .facet(col="Gender", row="User Type") .share(y=False) .add(so.Bars(), so.Hist(stat="proportion")) ) ``` ![The proportion of cycling duration by membership category.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_cycling_gender_membership.png) The proportion of cycling duration by membership category. Image by Author From the chart, we can see that for member users, regardless of gender, the distribution of cycling duration is more purposeful, tending to short-term cycling to quickly reach their destination. For ordinary users, users with 'Unknown' gender have a more casual cycling duration and longer cycling times. It seems that these users are there to temporarily get on their bikes and see the scenery? Therefore, in the fifth step, I want to see the distribution of bicycle usage times between stations to verify my guess. Since displaying so many stations on the chart can't be done, I first aggregate the sampled data by `Start Station ID` and `End Station ID` count. ```Python start_end_station = citibike_sample.groupby(["Start Station ID", "End Station ID"]).size().reset_index(name="Count") ``` Also, to avoid too many data points interfering with our analysis, I only took the data with the top 20% count for drawing. ```Python p8 = start_end_station["Count"].quantile(.8) start_end_filtered = start_end_station[start_end_station["Count"] >= p8] ``` Then use a scatter plot to plot the data and use the size of the point to represent the count size. ```Python ( so.Plot(start_end_filtered, x="Start Station ID", y="End Station ID", pointsize="Count", color="Count") .add(so.Dots()) ) ``` ![Distribution of rides between stations.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/so_scatter_plot_stations.png) Distribution of rides between stations. Image by Author The chart shows that the number of rides is mainly distributed between stations with ID values of 3180 and 3220. Compared with the table data, this area is concentrated for office workers. There is also a lot of data distribution in the Station ID between 3260 and 3280. By comparing the table data, we can see many parks and tourist attractions in this area. This confirms our guess: in addition to office workers who tend to ride shared bicycles on weekdays, many tourists are willing to use shared bikes to go out and see the scenery on weekends. Therefore, for this city's shared bicycle operation department, the operation strategy can not only discount on weekdays to attract members to ride more. They can also use new user registration gifts or promote more attractions in the app on weekends to encourage tourists or temporary users to become member users. --- ## Room for Growth: Current Limitations of Objects Interface After demonstrating how the Seaborn objects interface helps us quickly perform data analysis in actual projects, I would like to discuss some improvements the objects interface needs to make based on my experience. First, there needs to be more performance in the drawing. As shown in the above project, when I use the original dataset to draw, the speed is languid, and Seaborn doesn't use the calculation ability of `Numpy` or `Python Arrow`. Second, there needs to be more documentation. So many APIs I can not find the specific use of the introduction, and I can only slowly fumble. And the API design doesn’t feel very mature to me yet. For example, I believe `so.Stat` and `so.Move` should be placed in the Data Mapping phase, but currently, they are placed in the Presentation phase through the add method, which needs to be revised. Finally, the selection of charts needs to be more rich. I initially planned to use pie charts and map charts in the city bike-sharing project, but I couldn't find them. Although I could write an extension myself, that's a different story. Also, when I want to layout the charts more complexly, I need to use Matplotlib's `subplots` API and integrate it with the `on` method, which still needs to be fully encapsulated. Despite these shortcomings, I am confident about the future of Seaborn. I think the team's choice of declarative graphical syntax has made Seaborn easier and more flexible to use. I hope the Seaborn community will become more active in the near future. --- ## Conclusion In this article, I introduced the objects interface feature in Seaborn 0.12. By introducing the benefits of declarative graphic syntax, I let you understand why the Seaborn team chose to evolve in this way. Also, to cater to readers who need to become more familiar with Seaborn, I introduced the differences and similarities in API design philosophy between the original Seaborn and the objects interface version. By taking you through an actual project analysis of city bike-sharing usage, you've seen first-hand how the objects interface API is used and my expectations for it. Always remember, the goal of data visualization is not just to display data, but to uncover the stories behind the data. I hope you found this article helpful. Feel free to comment and participate in the discussion if you have any questions or new ideas. I'm more than happy to answer your questions. ### Efficient k-Nearest Neighbors (k-NN) Solutions with NumPy URL: https://www.dataleadsfuture.com/efficient-k-nearest-neighbors-k-nn-solutions-with-numpy/ Last updated: 2025-03-26T02:12:54.000Z ## Introduction I have a friend who is a city planner. One day, he was tasked with reassessing the location suitability of thousands of gas stations in the city, needing to find the positions of the k-nearest gas stations to each one. How can we find the nearest k stations with little time? This is a practical application scenario of the k-nearest neighbors problem. As such, he came to me for help, hoping I could provide a high-performance solution. So I write down this article and which will guide you on efficiently solving the [k-nearest neighbors](https://en.wikipedia.org/wiki/K-nearest%5Fneighbors%5Falgorithm?ref=dataleadsfuture.com) problem using NumPy. By comparing it with a Python iterative solution, we will demonstrate the powerful performance of NumPy. In this article, we will delve into utilizing advanced NumPy features, such as broadcasting, fancy indexing, and sorting, to implement a high-performance k-nearest neighbors algorithm. After reading this article, you will able to: - Understand the k-nearest neighbors problem and its practical application scenarios - Learn how to use the NumPy library to solve the k-nearest neighbors problem - Understand in-depth how features such as NumPy broadcasting, fancy indexing, and sorting play a role in the algorithm - Compare the performance of NumPy with a Python iterative solution, exploring why NumPy is superior Let’s delve into the high-performance world of NumPy together, exploring how we can solve the k-nearest neighbors problem more quickly and effectively using only NumPy. --- ## Geometric Principles of Solving the k-NN Problem Let’s review the gas station problem my friend faced from a geometric perspective. Assuming we place all the gas stations on a two-dimensional plane, the distance between two gas stations is actually the [Euclidean distance](https://en.wikipedia.org/wiki/Euclidean%5Fdistance?ref=dataleadsfuture.com) between two points on the plane. The solution formula is as follows: But how should the distance between any two points be calculated? We can imagine the two-dimensional plane as a chessboard, simplify the gas stations to six, and sequentially arrange these six points along the horizontal and vertical edges of the chessboard, as shown in the figure: ![Arrange these six points on the chessboard.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-77.png) Arrange these six points on the chessboard. Image by Author Then the grid where the extensions of any two points intersect can represent the distance between these two points. When i=j, the two points are the same, and the distance should be 0. Assuming that k=2 here, we only need to sort the distances from each point to other points in ascending order and take the points corresponding to the first three distances (including itself), which are the two closest other points to this point. ![After sorting, we can get the 3 points that are closest to each other.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-78.png) After sorting, we can get the 3 points that are closest to each other. Image by Author --- ## Traditional Python Iterative Solution As a performance benchmark, let’s first look at how the traditional Python iterative solution works. The idea of this solution is relatively simple: 1. To calculate the Euclidean distance between the coordinate point itself and other coordinate points in the list. 2. Then compare the distances from the current point to other points. 3. Take the top k points that meet the requirements. Next is the code part. First, we randomly generate six coordinate points. Since we will use the same coordinates as a comparison later, we need to add a `seed` to the `random` package. ```Python import random import matplotlib.pyplot as plt %matplotlib inline plt.style.use('seaborn-v0_8-whitegrid') random.seed(5) def generate_points(n: int=6) -> list[tuple]: points = [] for i in range(n): points.append((random.randint(0, 100), random.randint(0, 100))) return points ``` Next, start calculating the distance of each point to all points (including itself) in the list, which requires two iterations. ```Python def calc_dist(points: list[tuple]) -> list[list]: result = [] for i, left in enumerate(points): row = [left] for j, right in enumerate(points): dist = (left[0] - right[0])**2 + (left[1] - right[1])**2 row.append(dist) result.append(row) return result ``` Then, sort the distances between each point and other points and find the index of the point corresponding to the distance in the original list. ```Python def find_sorted_index(with_dist: list[list]) -> list[list]: results = [] for row in with_dist: dists = row[1:] sorted_dists = sorted(dists) indices = [dists.index(i) for i in sorted_dists] row[1:] = indices results.append(row) return results ``` The final return should be a two-dimensional array, where the first item in each row of the array is the current point, and the other items are the indexes of each point in the list after sorting the distance. Finally, we find each point that meets the conditions in the original coordinate list based on the index. ```Python def find_k_nearest(points: list[tuple], with_indices: list[list], k: int) -> list[tuple]: results = [] for row in with_indices: # Since the closest point to the current point is itself, we can get the point itself directly, so here is +2 k_indices = row[1:k+2] the_points = [points[i] for i in k_indices] results.append(the_points) return results ``` The result is a two-dimensional array, and each row of the array is the current point and the other two closest points. To facilitate our evaluation of the results, we use `Matplotlib` to draw all coordinate points and the lines from each coordinate to the two nearest coordinates. ```Python def draw_points(points: list[tuple]): x, y = [], [] for point in points: x.append(point[0]) y.append(point[1]) plt.scatter(x, y, s=100) def draw_lines(nearest: list[list]): for row in nearest: start = row[0] for end in row[1:]: plt.plot([start[0], end[0]], [start[1], end[1]], color='black') def orig_main(count: int = 6): k = 2 points = generate_points(count) with_dist = calc_dist(points) sorted_index = find_sorted_index(with_dist) nearest = find_k_nearest(points, sorted_index, k) return points, nearest points, nearest = orig_main(6) draw_points(points) draw_lines(nearest) ``` The result is as follows: ![Traditional Python Iterative Solution.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-79.png) Traditional Python Iterative Solution. Image by Author As you can see, six coordinates and corresponding lines have appeared on the chart. This chart will serve as a benchmark and will be compared with the results of using NumPy later to confirm the correctness of the algorithm. --- ## Basic Knowledge of Using NumPy Solution Next, let’s see how to solve this problem using NumPy. Before writing the code, we need to do some preheating on some basic concepts of NumPy. ### Broadcasting Since it involves placing a set of coordinate points on the chessboard horizontally (`shape=(1, 6)`) and vertically (`shape=(6, 1)`), and forming a `(6, 6)` matrix. After calculating the distance, it involves operations between two arrays of different sizes, so we need to use the broadcasting mechanism of NumPy. Here is an example: ```Python In: a = np.arange(6).reshape(1, 6) b = np.arange(6).reshape(6, 1) a + b Out: [[ 0 1 2 3 4 5] [ 1 2 3 4 5 6] [ 2 3 4 5 6 7] [ 3 4 5 6 7 8] [ 4 5 6 7 8 9] [ 5 6 7 8 9 10]] ``` As you can see, when a (1, 6) array and a (6, 1) array are added, the resulting shape is (6, 6). For the specific principles, please refer to the [official documentation](https://numpy.org/doc/stable/user/basics.broadcasting.html?ref=dataleadsfuture.com). The schematic diagram is as follows: ![How broadcasting works.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-80.png) How broadcasting works. Image by Author ### Sorting After solving the distance between any two points, we also need to sort the distances. Like the `sort()` function in the Python standard library, NumPy also has a function for sorting: `np.sort()`. Alternatively, the `ndarray.sort()` function can also be used for sorting. Since we are sorting the distances, we also need to find the index of each item in the original array after sorting. In NumPy, we can use `np.argsort()` to get it: ```Python In: x = np.array([2, 1, 4, 3, 5]) i = np.argsort(x) print(i) Out: [1 0 3 2 4] ``` Of course, we only need to focus on the k-nearest points, and we don’t need to know the order of distances. So we can use NumPy’s [argpartition()](https://numpy.org/doc/stable/reference/generated/numpy.argpartition.html?ref=dataleadsfuture.com) API, which can return the index of the smallest few points without sorting, which will perform better. ### Fancy Indexing In the traditional Python list, if we want to find a set of data by index, we need to iterate separately through the data list and index list, which has very poor performance. But NumPy provides fancy indexing to quickly find data corresponding to the index. Here is an example: ```Python In: x = np.array([8, 2, 4, 5, 3, 7, 1, 6]) ind = [0, 3, 7] print(x[ind]) Out: [8 5 6] ``` ![Fancy indexing can quickly find data corresponding to the index array.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-81.png) Fancy indexing can quickly find data corresponding to the index array. Image by Author Because fancy indexing is a set of integer arrays, there is a rule to follow: The data indexed reflects the shape of the broadcasted index array, which is unrelated to the shape of the data array. --- ## NumPy Solution After understanding some basics of NumPy, let’s see how to solve the k-NN problem using NumPy. Since here we are using a set of coordinate points to form an array, we need to use NumPy’s [structured\_array](https://numpy.org/doc/stable/user/basics.rec.html?ref=dataleadsfuture.com): ```Python import numpy as np from numpy import ndarray random.seed(5) def structured_array(points: list[tuple]) -> ndarray: dt = np.dtype([('x', 'int'), ('y', 'int')]) return np.array(points, dtype=dt) ``` Next, add an extra dimension to the original one-dimensional array in the horizontal and vertical directions, turning it into two sides of a two-dimensional chessboard: Then use the broadcasting mechanism to calculate the distance between each point. Finally, get a (6, 6) two-dimensional array: ```Python def np_find_dist(s_array: ndarray) -> ndarray: a = s_array.reshape(6, 1) b = s_array.reshape(1, 6) dist = (a['x'] - b['x'])**2 + (a['y'] - b['y'])**2 return dist ``` Then, use the `argpartition` method to find out the indexes of the two points with the smallest distance in each row: ```Python def np_k_nearest(dist: ndarray, k: int) -> ndarray: k_indices = np.argpartition(dist, k+1, axis=1)[:, :k+1] return k_indices ``` We still need two `Matplotlib `drawing methods to evaluate the correctness of the results: ```Python def np_draw_points(s_array: ndarray): plt.scatter(s_array['x'], s_array['y'], s=100) def np_draw_lines(s_array: ndarray, k_indices: ndarray, k: int): for i in range(s_array.shape[0]): for j in k_indices[i, :k+1]: plt.plot([s_array[i]['x'], s_array[j]['x']], [s_array[i]['y'], s_array[j]['y']], color='black') ``` Finally, write a main method to integrate all the code together: ```Python def np_main(count: int = 6): k = 2 points = generate_points(count) s_array = structured_array(points) np_dist = np_find_dist(s_array) k_indices = np_k_nearest(np_dist, k) results = [s_array[k_indices[i, :k+1]] for i in range(s_array.shape[0])] return results, s_array, k_indices, k results, s_array, k_indices, k = np_main(6) np_draw_points(s_array) np_draw_lines(s_array, k_indices, k) ``` Just looking at the code, it’s already much simpler than the Python iterative version. Next, we compare the results with the chart: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-82.png) The k-NN result of the NumPy solution. Image by Author See, the results are exactly the same! --- ## Performance Comparison of the Two Solutions Finally, let’s compare the execution performance of the two solutions. Here we still use `%timeit` for evaluation. First is the Python iterative way. Let’s see how long it takes to expand to 1,000 coordinates: ![The execution time of Python Iterative solution.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-83.png) The execution time of Python Iterative solution. Image by Author Then it’s the NumPy implementation. See how long it takes for 1,000 coordinates: ![The execution time of NumPy solution.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-84.png) The execution time of NumPy solution. Image by Author Surprised, right? The performance has improved hundreds of times, so my friend doesn’t have to worry about being unable to calculate it. --- ## Conclusion This article taught us how to use NumPy’s broadcasting, fancy indexing, and sorting to efficiently solve the k-nearest neighbors problem. We also compared the performance of NumPy with the Python iterative solution and deeply understood why NumPy can perform better in solving such problems. To recap, we learned the following: - The definition and practical application scenarios of the k-nearest neighbors problem - How to use the NumPy library to solve the k-nearest neighbors problem - The application of NumPy’s broadcasting, fancy indexing, sorting, and other features in algorithm implementation - The performance comparison analysis between NumPy and the Python brute force solution Although this article provides an efficient k-nearest neighbors solution, this is just a starting point. In future articles, I will reinterpret the solution to this problem using advanced algorithms and data structures, showing you more efficient and usable algorithm skills. Stay tuned for future articles. If you are interested in this article, feel free to comment, and I will answer them individually. ### Python Lists Vs. NumPy Arrays: A Deep Dive into Memory Layout and Performance Benefits URL: https://www.dataleadsfuture.com/python-lists-vs-numpy-arrays-a-deep-dive-into-memory-layout-and-performance-benefits/ Last updated: 2025-12-19T07:17:24.000Z In this article, we will delve into the memory design differences between native [Python lists](https://docs.python.org/3/tutorial/datastructures.html?ref=dataleadsfuture.com) and [NumPy arrays](https://numpy.org/doc/stable/reference/generated/numpy.array.html?ref=dataleadsfuture.com#numpy-array), revealing why NumPy can provide better performance in many cases. We will compare data structures, memory allocation, and access methods, showcasing the power of NumPy arrays. --- ## Introduction Imagine you are preparing to go to the library to find a book. Now, you discover that the library has two shelves: The first shelf is filled with various exquisite boxes, some containing CDs, some containing pictures, and others containing books. Only the name of the item is attached to the box. This represents native Python lists, where each element has its memory space and type information. However, this approach has a problem: many empty spaces in the boxes, wasting shelf space. Moreover, when you want to find a specific book, you must look inside each box, which takes extra time. Now let’s look at the second shelf. This time there are no boxes; books, CDs, and pictures are all compactly placed together according to their categories. This is NumPy arrays, which store data in memory in a continuous fashion, improving space utilization. Since the items are all grouped by category, you can quickly find a book without having to search through many boxes. This is why NumPy arrays are faster than native Python lists in many operations. --- ## Python Lists: A Flexible but Less Efficient Solution ### Everything in Python is an object Let’s start with the Python interpreter: although CPython is written in C, Python variables are not basic data types in C, but rather C structures that contain values and additional information. Take a Python integer `x = 10_000` as an example, `x` is not a basic type on the stack. Instead, it is a pointer to a memory heap object. If you delve into the source code of `Python 3.10`, you’ll find that the C structure that `x` points to is as shown in the following figure: ![Python integer vs. C native integer.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-69.png) Python integer vs. C native integer. Image by Author. The `PyObject_HEAD` contains information such as reference count, type information, and object size. ### Python lists are objects containing a series of objects From this, we can deduce that a Python list is also an object, except that it contains pointers to other objects. We can create a list that contains only integers: ```Python integer_list = [1, 2, 3, 4, 5] ``` We can also create a list that includes multiple object types: ```Python mixed_list = [1, "hello", 3.14, [1, 2, 3]] ``` ### Pros and cons of Python lists As we can see, Python lists contain a series of pointer objects. These pointers, in turn, point to other objects in memory. The advantage of this approach is flexibility. You can put any object in a Python list without worrying about type errors. However, the downside is also evident: ![Python lists contain a series of pointer objects.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-70.png) Python lists contain a series of pointer objects. Image by Author The objects pointed to by each pointer are scattered in memory. When you traverse a Python list, you need to look up the memory location of each object based on the pointer, resulting in lower performance. --- ## NumPy Arrays: A Contiguous Memory Layout for Enhanced Performance Next, let’s explore the components and arrangement of NumPy arrays, and how it benefits [cache locality](https://en.wikipedia.org/wiki/Locality%5Fof%5Freference?ref=dataleadsfuture.com) and [vectorization](https://en.wikipedia.org/wiki/Array%5Fprogramming?ref=dataleadsfuture.com). ### NumPy arrays: structure and memory layout According to NumPy’s internal description, NumPy arrays consist of two parts: 1. One part stores the metadata of the NumPy array, which describes the data type, array shape, etc. 2. The other part is the data buffer, which stores the values of array elements in a tightly packed arrangement in memory. ![NumPy arrays: structure and memory layout.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-71.png) NumPy arrays: structure and memory layout. Image by Author ### Memory layout of NumPy arrays When we observe the `.flags` attribute of a ndarray, we find that it includes: ```Python In 1: np_array = np.arange(6).reshape(2, 3, order='C') np_array.flags Out 1: C_CONTIGUOUS : True F_CONTIGUOUS : False OWNDATA : False WRITEABLE : True ALIGNED : True WRITEBACKIFCOPY : False ``` - `C_CONTIGUOUS`, which indicates whether the data can be read using row-major order. - `F_CONTIGUOUS`, which indicates whether the data can be read using column-major order. Row-major order is the data arrangement used by the C language, denoted by `order=’C’`. It means that data is stored row by row. Column-major order, on the other hand, is used by FORTRAN, denoted by `order=’F’`, and stores data column by column. ![Memory layout of NumPy arrays.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-72.png) Memory layout of NumPy arrays. Image by Author ### Advantages of NumPy’s memory layout Since `ndarray` is designed for matrix operations, all its data types are identical, with the same byte size and interpretation. This allows data to be tightly packed together, bringing advantages in cache locality and vectorized computation. --- ## Cache Locality: How NumPy’s Memory Layout Improves Cache Utilization ### What is CPU cache NumPy’s contiguous memory layout helps improve cache hit rates because it matches how CPU caches work. To better explain this, let’s first understand the basic concept of [CPU cache](https://en.wikipedia.org/wiki/CPU%5Fcache?ref=dataleadsfuture.com). A CPU cache is a small, high-speed storage area between the CPU and main memory (RAM). The purpose of the CPU cache is to speed up data access in memory. ![A CPU cache is a small, high-speed storage area between the CPU and main memory (RAM).](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-73.png) A CPU cache is a small, high-speed storage area between the CPU and main memory (RAM). Image by Author When the CPU needs to read or write data, it first checks if it is already in the cache. The CPU can read directly from the cache if the required data is in the cache (cache hit). If the data is not present (cache miss), the CPU loads the data from RAM and stores it in the cache for future use. CPU caches are usually organized in [cache lines](https://open-cas.github.io/cache%5Fline.html?ref=dataleadsfuture.com), which are contiguous memory addresses. When the CPU accesses RAM, the cache loads the entire cache line into the high-speed cache. This means that if the CPU accesses neighboring memory addresses, subsequent accesses are more likely to hit the cache after loading a cache line, thus improving performance. ![When the CPU accesses RAM, the cache loads the entire cache line into the high-speed cache.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-74.png) When the CPU accesses RAM, the cache loads the entire cache line into the high-speed cache. Image by Author ### How NumPy utilizes cache NumPy’s contiguous memory layout takes advantage of this fact. NumPy arrays store data in continuous memory addresses, which helps improve cache locality. When accessing an element in the array, the entire cache line (containing neighboring array elements) is loaded into the cache. As you traverse the array, you access each element in a sequence. Because array elements are stored contiguously in memory, cache hits are more likely during the traversal, improving performance. This is similar to going to the library to read a book. Instead of just grabbing the book you need, you also take out related books and place them on the table. This way, when you need to consult associated materials, they’re easily accessible and more efficient than searching the shelves. --- ## Vectorization: Unleashing the Power of NumPy’s Memory Layout ### What is vectorization Vectorization is a technique that leverages the [Single Instruction Multiple Data (SIMD)](https://en.wikipedia.org/wiki/Single%5Finstruction,%5Fmultiple%5Fdata?ref=dataleadsfuture.com) features of CPUs or GPUs to perform multiple data operations simultaneously. Vectorized operations can significantly improve code execution efficiency by simultaneously processing multiple data items. NumPy’s contiguous memory layout facilitates vectorized operations. ### Why is vectorization suitable Consider yourself a delivery person who must deliver packages to various households daily. Suppose the packages are arranged sequentially in the vehicle, and the houses are numbered along the street. In that case, the delivery person can efficiently deliver packages along the street in order. This efficient method is analogous to NumPy’s memory layout, which brings the following benefits in vectorization: - **Data alignment**: NumPy arrays’ contiguous memory layout ensures that data is aligned in memory in a vectorization-friendly manner. This allows the CPU to efficiently load and process data in NumPy. - **Sequential access pattern**: NumPy’s tightly packed data in memory helps improve vectorization performance. The sequential access pattern also takes full advantage of CPU cache and prefetching, reducing memory access latency. - **Simplified code**: NumPy offers a range of functions (e.g., `np.add`, `np.multiply`) and operations (e.g., array slicing) that automatically handle vectorized operations. You can write concise and efficient code without worrying about the underlying implementation. --- ## Copies and Views: NumPy’s Excellent Design for Performance Optimization In the earlier discussion, we discussed how NumPy leverages its contiguous memory layout to achieve performance advantages. Now, let’s discuss how NumPy gains performance benefits through copies and views. ### What are copies and views Copies and views are two options that define the relationship between existing data and the original array. Based on the characteristics of these two options, they can be summarized as follows: - **Copy**: Uses a different memory space than the original array, but the data content is the same. - **View**: References the same memory address as the original array. ![A copy may have multiple views.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-75.png) A copy may have multiple views. Image by Author If we compare this to a book, a view is like a bookmark placed in the book, without creating a copy of the book. On the other hand, a copy is a duplicate of the original book, containing a complete replica, including text and images. When you add notes to this copy, the original book remains unaffected. ### Making good use of both features Utilizing the characteristics of views and copies can help us write concise and efficient code. Let’s take arithmetic operations as an example. A new copy will be created if you use `a = a + 1`. However, if you use `a += 1` or `np.add`, broadcasting is applied, and the addition is performed directly on the original array. Consider the following code, which measures the execution time: ```Python import numpy as np import time def test(): x = np.ones(100_000_000, dtype='int8') y = np.ones(100_000_000, dtype='int8') start = time.monotonic() for _ in range(100): x = x * 4 + y * 3 x = np.ones(100_000_000, dtype='int8') print(f'x = x * 4 + y * 3: needs {time.monotonic() - start:.2f} seconds.') start = time.monotonic() for _ in range(100): x *= 4 x += y * 3 x = np.ones(100_000_000, dtype='int8') print(f'x *= 4; x += y * 3: needs {time.monotonic() - start:.2f} seconds.') start = time.monotonic() for _ in range(100): np.multiply(x, 4, out=x) np.multiply(y, 3, out=y) np.add(x, y, out=x) x = np.ones(100_000_000, dtype='int8') y = np.ones(100_000_000, dtype='int8') print(f'use functions needs {time.monotonic() - start:.2f} seconds.') if __name__ == "__main__": test() ``` Executing the above code will yield a result similar to the following: ![Using views for calculations takes less time.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-76.png) Using views for calculations takes less time. Screenshot by Author As we can see from the results, using views for calculations takes less time. ### Distinguishing between copies and views Confirming whether the result is a view or a copy every time a calculation is performed would require much effort. However, there’s a more straightforward way to verify this: - Use `may_share_memory` to determine if the two arrays in the argument refer to the same memory space. This judgment may not be strict. True doesn’t necessarily mean the arrays are shared, but False confirms that the arrays are definitely not shared. - If you need a more accurate answer, you can use the `share_memory` function. However, this function takes longer to execute than `may_share_memory`. --- ## Conclusion In summary, we have learned about the differences in memory arrangement between NumPy arrays and native Python lists. Due to the contiguous arrangement of the same data type in NumPy’s array, significant performance advantages are achieved in both Cache Locality and Vectorization. Separating views and copies in NumPy’s design provides greater flexibility for code execution performance and memory management. In the upcoming series of articles, I will start from the basics and reiterate the best practices for data science at work. If you have any suggestions or questions, please feel free to comment, and I will address them individually. ### Supercharge Your Python Asyncio With Aiomultiprocess: A Comprehensive Guide URL: https://www.dataleadsfuture.com/supercharge-your-python-asyncio-with-aiomultiprocess-a-comprehensive-guide/ Last updated: 2025-03-26T02:11:20.000Z In this article, I will take you into the world of [aiomultiprocess](https://aiomultiprocess.omnilib.dev/en/latest/index.html?ref=dataleadsfuture.com), a library that combines the powerful capabilities of Python [asyncio](https://docs.python.org/3/library/asyncio.html?ref=dataleadsfuture.com) and [multiprocessing](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com). This article will explain through rich code examples and best practices. By the end of this article, you will understand how to leverage the powerful features of aiomultiprocess to enhance your Python applications, just like a head chef leading a team of chefs to create a delicious feast. --- ## Introduction Imagine that you want to invite your colleagues over for a big meal on the weekend. How would you do it? As an experienced chef, you certainly wouldn’t cook one dish at a time; that would be too slow. You would efficiently use your time, letting multiple tasks happen simultaneously. For example, while you wait for the water to boil, you can step away to wash vegetables. This way, you can throw the vegetables into the pot when the water is boiling. This is the charm of concurrency. However, recipes can often be cruel: you need to keep stirring when making soup; vegetables need to be washed and chopped; you also need to bake bread, fry steaks, and more. When there are many dishes to prepare, you’ll be overwhelmed. Fortunately, your colleagues won’t just sit around waiting to eat. They will come into the kitchen to help you, with each additional person acting like an additional working process. This is the powerful combination of multiprocessing and concurrency. The same is true for code. Even with asyncio, has your Python application still encountered bottlenecks? Are you looking for ways to further improve the performance of your concurrent code? If so, `aiomultiprocess` is the answer you’ve been looking for. --- ## How to Install and Basic Usage ### Installation If you use pip, install it like this: ```Bash python -m pip install aiomultiprocess ``` If you use Anaconda, install it from conda-forge: ```Bash conda install -c conda-forge aiomultiprocess ``` ### Basic usage `aiomultiprocess` consists of three main classes: `Process` is the base class for the other two classes and is used to start a process and execute a coroutine function. You won’t usually need to use this class. `Worker` is used to start a process, execute a coroutine function, and return the result. We also won’t be using this class. `Pool` is the core class we will focus on. Like [multiprocessing.Pool](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#module-multiprocessing.pool), it starts a process pool, but its context needs to be managed using `async with`. We will use the two methods of Pool: `map` and `apply`. The `map` method accepts a coroutine function and an iterable. The `Pool` will iterate over the iterable and assign the coroutine function to run on various processes. The result of the map method can be asynchronously iterated using `async for:` ```Python import asyncio import random import aiomultiprocess async def coro_func(value: int) -> int: await asyncio.sleep(random.randint(1, 3)) return value * 2 async def main(): results = [] async with aiomultiprocess.Pool() as pool: async for result in pool.map(coro_func, [1, 2, 3]): results.append(result) print(results) if __name__ == "__main__": asyncio.run(main()) ``` The `apply` method accepts a coroutine function and the required argument tuple for the function. According to the scheduler’s rules, the `Pool` will assign the coroutine function to an appropriate process for execution. ```Python import asyncio import random import aiomultiprocess async def coro_func(value: int) -> int: await asyncio.sleep(random.randint(1, 3)) return value * 2 async def main(): tasks = [] async with aiomultiprocess.Pool() as pool: tasks.append(pool.apply(coro_func, (1,))) tasks.append(pool.apply(coro_func, (2,))) tasks.append(pool.apply(coro_func, (3,))) results = await asyncio.gather(*tasks) print(results) # Output: [2, 4, 6] if __name__ == "__main__": asyncio.run(main()) ``` --- ## Implementation Principle and Practical Examples ### Implementation principle of aiomultiprocess.Pool [In a previous article](https://www.dataleadsfuture.com/harnessing-multi-core-power-with-asyncio-in-python/), I explained how to distribute asyncio tasks across multiple CPU cores. The general approach is to start a process pool in the main process using [loop.run\_in\_executor](https://docs.python.org/3/library/asyncio-eventloop.html?ref=dataleadsfuture.com#asyncio.loop.run%5Fin%5Fexecutor). Then, an [asyncio event loop](https://docs.python.org/3/library/asyncio-eventloop.html?ref=dataleadsfuture.com) is created in each process in the process pool, and the coroutine functions are executed in their respective loops. The schematic is as follows: ![This diagram shows the way to integrate asyncio and multiprocessing.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-65.png) This diagram shows the way to integrate asyncio and multiprocessing. Image by Author The implementation of `aiomultiprocess.Pool` is similar. It includes `scheduler`, `queue`, and `process` as its three components. - The `scheduler` can be understood as the head chef, responsible for allocating tasks in a suitable way to each chef. Of course, you can hire (implement) a head chef suitable for your needs. - The `queue` is like the kitchen’s assembly line. Strictly speaking, it includes an order line and a delivery line. The head chef passes the menu through the order line to the chefs, and the chefs return the completed dishes through the delivery line. - The `process` is like the chefs in the restaurant. They each handle several dishes concurrently according to the allocation. Each time a dish is ready, it will be handed over in the allocated order. The entire schematic is shown below: ![Aiomultiprocess consists of three components: scheduler, queue, and process.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-66.png) Aiomultiprocess consists of three components: scheduler, queue, and process. Image by Author --- ## Real-world Example Based on the introduction provided earlier, you should now understand how to use `aiomultiprocess`. Let’s dive into a real-world example to experience the power of it. Or, you can find a free version here: [Aiomultiprocess: Super Easy Integrate Multiprocessing & Asyncio in PythonEven no need to know much about asyncio and multiprocessing![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_QnAEw78kRY7HSi3PAZLPhA.webp)](https://www.dataleadsfuture.com/aiomultiprocess-super-easy-integrate-multiprocessing-asyncio-in-python/) First, we’ll use a remote call and a loop calculation to simulate the process of data retrieval and processing in real life. This method demonstrates that IO-bound and CPU-bound tasks are often mixed together, and the boundary between them is not so clear-cut. ```Python import asyncio import random import time from aiohttp import ClientSession from aiomultiprocess import Pool def cpu_bound(n: int) -> int: result = 0 for i in range(n*100_000): result += 1 return result async def invoke_remote(url: str) -> int: await asyncio.sleep(random.uniform(0.2, 0.7)) async with ClientSession() as session: async with session.get(url) as response: status = response.status result = cpu_bound(status) return result ``` Next, let’s use the traditional asyncio approach to call this task 30 times as a baseline: ```Python async def main(): start = time.monotonic() tasks = [asyncio.create_task(invoke_remote("https://www.example.com")) for _ in range(30)] await asyncio.gather(*tasks) print(f"All jobs done in {time.monotonic() - start} seconds") if __name__ == "__main__": asyncio.run(main()) ``` ![The code is run using the traditional asyncio method.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-67.png) The code is run using the traditional asyncio method. Screenshot by Author The code execution results are shown in the figure, and it takes approximately 21 seconds. Now let’s see how much aiomultiprocess can improve this. Using aiomultiprocess is simple. The original concurrent code does not need to be modified. You only need to adjust the code in the main method to run inside the Pool: ```Python async def main(): start = time.monotonic() async with Pool() as pool: tasks = [pool.apply(invoke_remote, ("https://www.example.com",)) for _ in range(30)] await asyncio.gather(*tasks) print(f"All jobs done in {time.monotonic() - start} seconds") if __name__ == "__main__": asyncio.run(main()) ``` ![Simply use the modified version of aiomultiprocess.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-68.png) Simply use the modified version of aiomultiprocess. Screenshot by Author As you can see, the code using aiomultiprocess takes only 14 seconds to complete on my laptop. The performance improvement would be even greater on a more powerful computer. --- ## Detailed Best Practices Finally, based on my experience, let me share some more practical best practices. ### Use pool only Although `aiomultiprocess` also provides the `Process` and `Worker` classes for us to choose from, we should always use the `Pool` class to ensure maximum efficiency due to the significant resource consumption of creating processes. ### How to use queues [In a previous article](https://www.dataleadsfuture.com/unleashing-the-power-of-python-asyncios-queue/), I explained how to use `asyncio.Queue` to implement the producer-consumer pattern to balance resources and performance. In `aiomultiprocess`, we can also use queues. However, since we are in a process pool, we cannot use `asyncio.Queue`. At the same time, we cannot directly use `multiprocessing.Queue` in the process pool. In this case, you should use `multiprocessing.Manager().Queue()` to create a queue, with the code as follows: ```Python import random import asyncio from multiprocessing import Manager from multiprocessing.queues import Queue from aiomultiprocess import Pool async def worker(name: str, queue: Queue): while True: item = queue.get() if not item: print(f"worker: {name} got the end signal, and will stop running.") queue.put(item) break await asyncio.sleep(random.uniform(0.2, 0.7)) print(f"worker: {name} begin to process value {item}", flush=True) async def producer(queue: Queue): for i in range(20): await asyncio.sleep(random.uniform(0.2, 0.7)) queue.put(random.randint(1, 3)) queue.put(None) async def main(): queue: Queue = Manager().Queue() producer_task = asyncio.create_task(producer(queue)) async with Pool() as pool: c_tasks = [pool.apply(worker, args=(f"worker-{i}", queue)) for i in range(5)] await asyncio.gather(*c_tasks) await producer_task if __name__ == "__main__": asyncio.run(main()) ``` ### Using initializer to initialize resources Suppose you need to use an `aiohttp` session or a database connection pool in a coroutine method, but we cannot pass arguments when creating tasks in the main process because these objects cannot be pickled. An alternative is to define a global object and an initialization method. In this initialization method, access the global object and perform initialization. Just like [multiprocessing.Pool](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#module-multiprocessing.pool), `aiomultiprocess.Pool` can accept an initialization method and corresponding initialization parameters when initialized. This method will be called to complete the initialization when each process starts: ```Python import asyncio from aiomultiprocess import Pool import aiohttp from aiohttp import ClientSession, ClientTimeout session: ClientSession | None = None def init_session(timeout: ClientTimeout = None): global session session = aiohttp.ClientSession(timeout=timeout) async def get_status(url: str) -> int: global session async with session.get(url) as response: status_code = response.status return status_code async def main(): url = "https://httpbin.org/get" timeout = ClientTimeout(2) async with Pool(initializer=init_session, initargs=(timeout,)) as pool: tasks = [asyncio.create_task(pool.apply(get_status, (url,))) for i in range(3)] status = await asyncio.gather(*tasks) print(status) if __name__ == "__main__": asyncio.run(main()) ``` ### Exception handling and retries Although `aiomultiprocess.Pool` provides the `exception_handler` parameter to help with exception handling, if you need more flexibility, you need to combine it with `asyncio.wait`. For the usage of `asyncio.wait`, you can refer to [my previous article](https://www.dataleadsfuture.com/use-these-methods-to-make-your-python-concurrent-tasks-perform-better/). With `asyncio.wait`, you can get tasks that encounter exceptions. After extracting the task, you can make some adjustments and then re-execute the task, as shown in the code below: ```Python import asyncio import random from aiomultiprocess import Pool async def worker(): await asyncio.sleep(0.2) result = random.random() if result > 0.5: print("will raise an exception") raise Exception("something error") return result async def main(): pending, results = set(), [] async with Pool() as pool: for i in range(7): pending.add(asyncio.create_task(pool.apply(worker))) while len(pending) > 0: done, pending = await asyncio.wait(pending, return_when=asyncio.FIRST_EXCEPTION) print(f"now the count of done, pending is {len(done)}, {len(pending)}") for result in done: if result.exception(): pending.add(asyncio.create_task(pool.apply(worker))) else: results.append(await result) print(results) if __name__ == "__main__": asyncio.run(main()) ``` ### Using Tenacity for retries Of course, we have more flexible and powerful options for exception handling and retries, such as using the `Tenacity` library, which I explained in [this article](https://www.dataleadsfuture.com/conquer-retries-in-python-using-tenacity-an-end-to-end-tutorial/). With `Tenacity`, the code above can be significantly simplified. You just need to add a decorator to the coroutine method, and the method will automatically retry when an exception is thrown. ```Python import asyncio from random import random from aiomultiprocess import Pool from tenacity import * @retry() async def worker(name: str): await asyncio.sleep(0.3) result = random() if result > 0.6: print(f"{name} will raise an exception") raise Exception("something wrong") return result async def main(): async with Pool() as pool: tasks = pool.map(worker, [f"worker-{i}" for i in range(5)]) results = await tasks print(results) if __name__ == "__main__": asyncio.run(main()) ``` ### Using tqdm to indicate progress I like `tqdm` because it can always tell me how far the code has run when I’m waiting in front of the screen. [This article](https://www.dataleadsfuture.com/using-tqdm-with-asyncio-in-python/) also explains how to use it. Since `aiomultiprocess` uses asyncio’s API to wait for tasks to complete, it is also compatible with `tqdm`: ```Python import asyncio from random import uniform from aiomultiprocess import Pool from tqdm.asyncio import tqdm_asyncio async def worker(): delay = uniform(0.5, 5) await asyncio.sleep(delay) return delay * 10 async def main(): async with Pool() as pool: tasks = [asyncio.create_task(pool.apply(worker)) for _ in range(1000)] results = await tqdm_asyncio.gather(*tasks) print(results[:10]) if __name__ == "__main__": asyncio.run(main()) ``` --- ## Conclusion Running asyncio code is like a chef cooking a meal. Even if you can improve efficiency by running different tasks concurrently, you’ll eventually encounter bottlenecks. The simplest solution at this point is to add more chefs to increase the parallelism of the cooking process. `Aiomultiprocess` is such a powerful Python library. By allowing concurrent tasks to run on multiple processes, it perfectly breaks through the performance bottlenecks caused by asyncio’s single-threaded nature. The use and best practices of `aiomultiprocess` in this article are based on my work experience. If you’re interested in any aspect, feel free to comment and join the discussion. ### Conquer Retries in Python Using Tenacity: An End-to-End Tutorial URL: https://www.dataleadsfuture.com/conquer-retries-in-python-using-tenacity-an-end-to-end-tutorial/ Last updated: 2026-01-23T06:37:10.000Z This article will discuss Tenacity’s basic usage and customization capabilities. This instrumental Python library provides a retry mechanism. We will also explore Tenacity’s retry and exception-handling capabilities through a practical example. --- ## Introduction Imagine you’re managing hundreds of web services, some located overseas (with high latency) and others pretty old (and not very stable). How would you feel? My colleague Wang is in such a predicament. He told me that he was pretty frustrated: Every day, he needs to check the invocation status of these remote services, and he often encounters timeout issues or other anomalies. Troubleshooting is particularly challenging. Moreover, much of the client-side code was written by his predecessors, making it challenging to perform substantial refactoring. So, the services have to continue running as they are. It would be great if there was a way to automatically reconnect these remote calls after an exception occurs. With tears in his eyes, Wang looked at me. I assured him it was no problem and introduced him to a new tool from my toolbox: [Tenacity](https://github.com/jd/tenacity?ref=dataleadsfuture.com). With just one decorator, the existing code can gain retry capabilities. Let’s see how to use it. --- ## Installation and Basic Usage Since Tenacity’s official website only offers a simple API document, let’s start with the library’s installation and some basic usage. ### Installation If you’re using pip, simply run the following: ```Bash python -m pip install tenacity ``` If you’re using Anaconda, Tenacity is not in the default channel, so you need to install it from `conda-forge`: ```Bash conda install -c conda-forge tenacity ``` ### Basic usage After installing Tenacity, let’s look at some basic usage of the library. Simply add an `@retry` decorator and your code will have retry capabilities: ```Python @retry() async def coro_func(): pass ``` If you want your code to stop retrying after a certain number of attempts, you can write it like this: ```Python @retry(stop=stop_after_attempt(5)) async def coro_func(): pass ``` Of course, to avoid frequent retries that may exhaust connection pools, I recommend adding a waiting time before each retry. For example, if you want to wait for 2 seconds before each connection: ```Python @retry(wait=wait_fixed(2)) async def coro_func(): pass ``` Although it’s not mentioned in the documentation, I prefer to wait an extra second longer than the last time before each retry to minimize resource waste: ```Python @retry(wait=wait_incrementing(start=1, increment=1, max=5)) async def coro_func(): pass ``` Finally, if the retry is caused by an `exception` being thrown in the method, it is best to throw the `exception` back out. This allows for more flexible exception handling when calling the method: ```Python @retry(reraise=True, stop=stop_after_attempt(3)) async def coro_func(): pass ``` --- ## Advanced Features: Custom Callbacks In addition to some common use cases, you may add your own retry determination logic, such as deciding based on the result of the method execution or printing the method invocation parameters before execution. In this case, we can use `Custom Callbacks` for customization. There are two ways to extend `Custom Callbacks`: One is the recommended approach from the documentation: writing an extension method. This method will be passed a `RetryCallState` instance as a parameter when executed. Through this parameter, we can obtain the wrapped method, the parameters of the method call, the returned result, and any thrown exceptions. For example, we can use this approach to judge the return value of a method and retry if the value is even: ```Python from tenacity import * def check_is_even(retry_state: RetryCallState): if retry_state.outcome.exception(): return True return retry_state.outcome.result() % 2 == 0 ``` Of course, before making this judgment, if an `exception` is thrown, retry directly. If you need to pass additional parameters in the extension method, you can add a wrapper outside the extension method. For example, this wrapper will pass a `logger` parameter. When the number of retries exceeds two, it will print the retry time, method name, and method parameters to the log: ```Python def my_before_log(logger: Logger): def my_log(retry_state: RetryCallState): fn = retry_state.fn args = retry_state.args attempt = retry_state.attempt_number if attempt > 2: logger.warning(f"Start retry method {fn.__name__} with args: {args}") return my_log ``` --- ## Real-World Network Example Finally, to give you a deep impression of using `Tenacity` in your projects, I will use a remote client project as an example to demonstrate how to integrate Tenacity’s powerful capabilities. This project will simulate accessing an HTTP service and deciding whether or not to retry based on the returned `status code`. Of course, to avoid wasting server resources due to long connection wait times, I will also add a 2-second timeout for each request. If a timeout occurs, the connection will be retried. ![Flow chart of the project.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-63.png) Flow chart of the project. Image by Author Before starting the code, I will implement several extension methods. One of the methods is to judge when a method’s retry count exceeds two, and print a warning message in the log: ```Python import asyncio import logging import random import sys import aiohttp from aiohttp import ClientTimeout, ClientSession from tenacity import * def ready_logger(stream, level) -> logging.Logger: logger = logging.getLogger(__name__) logger.setLevel(level) handler = logging.StreamHandler(stream) formatter = logging.Formatter('%(asctime)s - %(name)s - %(levelname)s - %(message)s') handler.setFormatter(formatter) logger.addHandler(handler) return logger logger = ready_logger(stream=sys.stderr, level=logging.DEBUG) def my_before_log(logger: logging.Logger): def log_it(retry_state: RetryCallState): fn = retry_state.fn attempt = retry_state.attempt_number if attempt > 2: logger.warning(f"Retrying method {fn.__name__} at the {attempt} attempt") return log_it ``` Another extension method is to judge the returned `status code`. If the status code is greater than 300, retry. Of course, timeouts will also trigger retries. ```Python def check_status(retry_state: RetryCallState) -> bool: outcome: Future = retry_state.outcome if outcome.exception(): return True return outcome.result() > 300 ``` Next, we have the implementation of the remote call method. After writing the method, remember to add Tenacity’s retry decorator. The strategy I use here is to retry up to 20 times, waiting for 1 second longer than the previous retry before each retry. Of course, don’t forget to add the two extension methods we just implemented: ```Python @retry(stop=stop_after_attempt(20), wait=wait_incrementing(start=1, increment=1, max=5), before=my_before_log(logger), retry=check_status) async def get_status(url_template: str, session: ClientSession) -> int: status_list = [200, 300, 400, 500] url = url_template.format(codes=random.choice(status_list)) print(f"Begin to get status from {url}") async with session.get(url) as response: return response.status async def main(): timeout: ClientTimeout = aiohttp.ClientTimeout(2) async with aiohttp.ClientSession(timeout=timeout) as session: tasks = [asyncio.create_task( get_status('https://httpbin.org/status/{codes}', session)) for _ in range(5)] result = await asyncio.gather(*tasks) print(result) if __name__ == "__main__": asyncio.run(main()) ``` ![After several retries, I finally got the correct result.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-64.png) After several retries, I finally got the correct result. Screenshot by Author Mission complete! Wasn’t that super simple? --- ## Conclusion I’m glad I helped Wang solve another problem. By using `Tenacity`, we can easily equip existing code with various retry mechanisms, thereby enhancing the robustness and self-recovery capabilities of the program. I would be even happier if this library could help you solve problems. Feel free to leave a comment and discuss. ### Introducing Python’s Parse: The Ultimate Alternative to Regular Expressions URL: https://www.dataleadsfuture.com/introducing-pythons-parse-the-ultimate-alternative-to-regular-expressions/ Last updated: 2026-07-09T07:36:07.000Z This article introduces a Python library called [parse](https://pypi.org/project/parse/?ref=dataleadsfuture.com) for quickly and conveniently parsing and extracting data from text, serving as a great alternative to [Python regular expressions](https://docs.python.org/3/library/re.html?ref=dataleadsfuture.com). And which covers the best practices with the [parse](https://pypi.org/project/parse/?ref=dataleadsfuture.com) library and a real-world example of parsing [nginx log text](http://nginx.org/en/docs/http/ngx%5Fhttp%5Flog%5Fmodule.html?ref=dataleadsfuture.com#log%5Fformat). --- ## Introduction I have a colleague named Wang. One day, he came to me with a worried expression, saying he encountered a complex problem: his boss wanted him to analyze the server logs from the past month and provide statistics on visitor traffic. I told him it was simple. Just use regular expressions. For example, to analyze nginx logs, use the following regular expression, and it’s elementary. ```Bash content: 192.168.0.2 - - [04/Jan/2019:16:06:38 +0800] "GET http://example.aliyundoc.com/_astats?application=&inf.name=eth0 HTTP/1.1" 200 273932 regular expression: (?\d+\.\d+\.\d+\.\d+)( - - \[)(?[\s\S]+)(?\][\s"]+)(?[A-Z]+) (?[\S]*) (?[\S]+)["] (?\d+) (?\d+) ``` But Wang was still worried, saying that learning regular expressions is too tricky. Although there are many ready-made examples online to learn from, he needs help with parsing uncommon text formats. Moreover, even if he could solve the problem this time, what if his boss asked for changes in the parsing rules when he submitted the analysis? Wouldn’t he need to fumble around for a long time again? Is there a simpler and more convenient method? I thought about it and said, of course, there is. Let’s introduce our protagonist today: the Python `parse` library. --- ## Installation & Setup As described on [the parse GitHub page](https://github.com/r1chardj0n3s/parse?ref=dataleadsfuture.com), it uses [Python’s format() syntax](https://docs.python.org/3/library/string.html?ref=dataleadsfuture.com#format-string-syntax) to parse text, essentially serving as a reverse operation of [Python f-strings](https://docs.python.org/3/reference/lexical%5Fanalysis.html?ref=dataleadsfuture.com#f-strings). Before starting to use `parse`, let’s see how to install the library. Direct installation with pip: ```Bash python -m pip install parse ``` Installation with conda can be more troublesome, as parse is not in the default conda channel and needs to be installed through conda-forge: ```Bash conda install -c conda-forge parse ``` After installation, you can use `from parse import *` in your code to use the library’s methods directly. --- ## Features & Usage The `parse` API is similar to [Python Regular Expressions](https://docs.python.org/3/library/re.html?ref=dataleadsfuture.com#functions), mainly consisting of the `parse`, `search`, and `findall` methods. Basic usage can be learned from [the parse documentation](https://pypi.org/project/parse/?ref=dataleadsfuture.com). ### Pattern format The parse format is very similar to the Python format syntax. You can capture matched text using `{}` or `{field_name}`. For example, in the following text, if I want to get the profile URL and username, I can write it like this: ```Bash content: Hello everyone, my Medium profile url is https://qtalen.medium.com, and my username is @qtalen. parse pattern: Hello everyone, my Medium profile url is {profile}, and my username is {username}. ``` Or you want to extract multiple phone numbers. Still, the phone numbers have different formats of country codes in front, and the phone numbers are of a fixed length of 11 digits. You can write it like this: ```Bash compiler = Parser("{country_code}{phone:11.11},") content = "0085212345678901, +85212345678902, (852)12345678903," results = compiler.findall(content) for result in results: print(result) ``` Or if you need to process a piece of text in an HTML tag, but the text is preceded and followed by an indefinite length of whitespace, you can write it like this: ```Bash content:
Hello World
pattern:
{:^}
``` In the code above, `{:11}` refers to the width, which means to capture at least 11 characters, equivalent to the regular expression `(.{11,})?`. `{:.11}` refers to the precision, which means to capture at most 11 characters, equivalent to the regular expression `(.{,11})?`. So when combined, it means `(.{11, 11})?`. The result is: ![Capture fixed-width characters.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-59.png) Capture fixed-width characters. Image by Author The most powerful feature of parse is its handling of time text, which can be directly parsed into Python datetime objects. For example, if we want to parse the time in an HTTP log: ```Bash content: [04/Jan/2019:16:06:38 +0800] pattern: [{:th}] ``` ### Retrieving results There are two ways to retrieve the results: 1. For capturing methods that use `{}` without a field name, you can directly use `result.fixed` to get the result as a tuple. 2. For capturing methods that use `{field_name}`, you can use `result.named` to get the result as a dictionary. ### Custom Type Conversions Although using `{field_name}` is already quite simple, the source code reveals that `{field_name}` is internally converted to `(?P.+?)`. So, `parse` still uses regular expressions for matching. `.+?` represents one or more random characters in non-greedy mode. ![The transformation process of parse format to regular expressions.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-60.png) The transformation process of parse format to regular expressions. Image by Author However, often we hope to match more precisely. For example, the text “my email is [xxx@xxx.com](mailto:xxx@xxx.com)”, `“my email is {email}”` can capture the email. Sometimes we may get dirty data, for example, “my email is xxxx@xxxx”, and we don’t want to grab it. Is there a way to use regular expressions for more accurate matching? That’s when the `with_pattern` decorator comes in handy. For example, for capturing email addresses, we can write it like this: ```Python @with_pattern(r'\b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b') def email(text: str) -> str: return text compiler = Parser("my email address is {email:Email}", dict(Email=email)) legal_result = compiler.parse("my email address is xx@xxx.com") # legal email illegal_result = compiler.parse("my email address is xx@xx") # illegal email ``` Using the `with_pattern` decorator, we can define a custom field type, in this case, `Email`which will match the email address in the text. We can also use this approach to match other complicated patterns. Want to add enterprise-level AI development projects to your resume? Check out the [****Generative AI Software Engineering Specialization**](https://imp.i384100.net/n4AoJM?ref=dataleadsfuture.com), where you'll pick up everything from the theory to hands-on AI software engineering. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [👉 Start your free trial now → ](https://imp.i384100.net/n4AoJM?ref=dataleadsfuture.com) --- ## A Real-world Example: Parsing Nginx Log After understanding the basic usage of parse, let’s return to the troubles of Wang mentioned at the beginning of the article. Let’s see how to parse logs if we have server log files for the past month. **Note:** We chose [NASA’s HTTP log dataset](https://ita.ee.lbl.gov/html/contrib/NASA-HTTP.html?ref=dataleadsfuture.com) for this experiment, which is free to use. The text fragment to be parsed looks like this: ![What is the text fragment look like.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-61.png) What is the text fragment look like. Screenshot by Author First, we need to preprocess the parse expression. This way, when parsing large files, we don’t have to compile the regular expression for each line of text, thus improving performance. ```Python from parse import Parser, with_pattern import pandas as pd # https://ita.ee.lbl.gov/html/contrib/NASA-HTTP.html FILE_NAME = "../../data/access_log_Jul95_min" compiler = Parser('{source} - - [{timestamp:th}] "{method} {path} {version}" {status_code} {length}\n') ``` Next, the `parse_line` method is the core of this example. It uses the preprocessed expression to parse the text, returning the corresponding match if there is one and an empty dictionary if not. ```Python def process_line(text: str) -> dict: parse_result = compiler.parse(text) return parse_result.named if parse_result else {} ``` Then, we use the `read_file` method to process the text line by line using a generator, which can minimize memory usage. However, due to the disk’s 4k capability limitations, this method may not guarantee performance. ```Python def read_file(name: str) -> list[dict]: result = [] with open(name, 'r') as f: for line in f: obj: dict = process_line(line) result.append(obj) return result ``` Since we need to perform statistics on the log files, we must use the `from_records` method to construct a `DataFrame` from the matched results. ```Python def build_dataframe(records: list[dict]) -> pd.DataFrame: result: pd.DataFrame = pd.DataFrame.from_records(records, index='timestamp') return result ``` Finally, in the `main` method, we put all the methods together and try to count the different `status_code` occurrences: ```Python def main(): records: list[dict] = read_file(FILE_NAME) dataframe = build_dataframe(records) print(dataframe.groupby('status_code').count()) ``` ![Wang’s troubles have been easily solved.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-62.png) Wang’s troubles have been easily solved. Image by Author That’s it. Wang’s troubles have been easily solved. --- ## Best Practices with parse Library Although the `parse` library is so simple that I only have a little to write about in the article. There are still some best practices to follow, just like regular expressions. ### Readability and maintainability To efficiently capture text and maintain expressions, it is recommended to always use `{field_name}` instead of `{}`. This way, you can directly use `result.named` to obtain key-value results. Using `Parser(pattern)` to preprocess the expression is recommended, rather than `parse(pattern, text)`. On the one hand, this can improve performance. On the other hand, when using `Custom Type Conversions`, you can keep the `pattern` and `extra_type` together, making it easier to maintain. ### Optimizing performance for large datasets If you look at the source code, you can see that `{}` and `{field_name}` use the regular expressions `(.+?)` and `(?P.+?)` for capture, respectively. Both expressions use the [non-greedy mode](https://docs.python.org/3/library/re.html?ref=dataleadsfuture.com#regular-expression-syntax). So when you use `with_pattern` to write your own expressions, also try to use non-greedy mode. At the same time, when writing `with_pattern`, if you use `()` for capture grouping, please use `regex_group_count` to specify the specific groups like this: [@with\_pattern](http://twitter.com/with%5Fpattern?ref=dataleadsfuture.com)`(r’((\d+))’, regex_group_count=2)` . Finally, if a group is not needed in with\_pattern, use `(?:x)` instead. `@with_pattern(r’(?:)(.*?)(?:)’, regex_group_count=1)` means you want to capture the content between input tags. The input tags will not be captured. --- ## Conclusion In this article, I changed my usual way of writing lengthy papers. By solving a colleague’s problem, I briefly introduced the use of the `parse` library. I hope you like this style. This article does not cover the detailed usage methods on the official website. Still, it introduces some best practices and performance optimization solutions based on my experience. At the same time, I explained in detail the use of the `parse` library to parse nginx logs with a practical example. As the new series title suggests, besides improving code execution speed and performance, using various tools to improve work efficiency is also a performance enhancement. This article helps data scientists simplify text parsing and spend time on more critical tasks. If you have any thoughts on this article, feel free to leave a comment and discuss. ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. --- next step Mr. Qian's Recommendation: Knowing a few AI coding tricks won't really set you apart at work. What you actually need is the ability to integrate AI-generated code into real enterprise codebases. I highly recommend the [****Generative AI Software Engineering Specialization**](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) by Vanderbilt University. It'll take you from being a casual user to becoming a truly AI-driven software engineer. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [👉 Start Learning for Free Now ](https://imp.i384100.net/qWVbNg?ref=dataleadsfuture.com) --- ## Further reading Learn how I used OpenCode and some open-source plugins to build an open-source coding workflow that can go toe-to-toe with Claude Code: [How I Use OpenCode, Oh-My-OpenCode-Slim, and OpenSpec to Build My Own AI Coding EnvironmentRide the wave of AI coding, don’t get swept away by it![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-a188d787-8df0-4038-a0af-b3331a7239bb.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/opencode_cover_3-1-07ed36c8-6dc8-41b5-b20d-7f165b4363e1.webp)](https://www.dataleadsfuture.com/how-i-use-opencode-oh-my-opencode-slim-and-openspec-to-build-my-own-ai-coding-environment/) By adding a reflective agent to the OpenSpec workflow, I managed to get DeepSeek-V4-Pro to perform at the level of Opus: [Reflection SDD: Use a Reflection Harness to Level Up Your OpenSpec WorkflowStop letting bad spec files tank your code quality![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-55cb8c76-9f19-4a00-b76a-16fb1b00d37d.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/9f04a900-aa9a-4fd2-bd70-6f3fe3783129-9e31999f-c521-4961-8cd6-917385524a29.webp)](https://www.dataleadsfuture.com/reflection-sdd-use-a-reflection-harness-to-level-up-your-openspec-workflow/) Mr. Qian shares his personal journey and walks you through how he went from having no relevant skills to landing a senior data scientist role at a major investment bank: [How to Become a Data Scientist If You Lacking the Necessary SkillsEmbrace Change, Love Learning, and Persist![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/icon/color_192_192-335a6a2d-fd42-4c20-8b73-abf1c318396b.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/thumbnail/Bruce_-cb6ef4c4-a747-4ee4-9896-d5a84768137a.webp)](https://www.dataleadsfuture.com/how-to-become-a-data-scientist-if-you-lacking-the-necessary-skills/) ### Unleashing the Power of Python Asyncio’s Queue URL: https://www.dataleadsfuture.com/unleashing-the-power-of-python-asyncios-queue/ Last updated: 2025-12-19T07:15:57.000Z In this article, I will explain the API usage and application scenarios of various queues in Python asyncio relaxedly. At the end of the article, I will demonstrate the practical usage of `asyncio.Queue` in a classic shopping scenario. --- ## Introduction ### Why do we need [asyncio.Queue](https://docs.python.org/3/library/asyncio-queue.html?ref=dataleadsfuture.com#queue) As readers who have read my previous articles know, I love asyncio because it is an almost perfect solution for concurrent programming. However, in a large-scale, highly concurrent project, a large number of uncontrollable concurrent tasks waiting will occupy system resources, leading to poor performance. Therefore, it is necessary to control the number of concurrent tasks. ### Why can’t we use [asyncio.Semaphore](https://docs.python.org/zh-cn/3/library/asyncio-sync.html?ref=dataleadsfuture.com#asyncio.Semaphore) In my previous article on synchronization primitives, I introduced using `Semaphore` locks to control the number of concurrent tasks running simultaneously. [Mastering Synchronization Primitives in Python Asyncio: A Comprehensive GuideBest practices for asyncio.Lock, asyncio.Semaphore, asyncio.Event and asyncio.Condition![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/Mastering_Synchronization_Primitives.png)](https://www.dataleadsfuture.com/mastering-synchronization-primitives-in-python-asyncio-a-comprehensive-guide/) Set the number of `Semaphore` locks first, and the tasks that get the lock will be executed, while those who don’t will wait. However, `asyncio.Semaphore` can only limit the concurrency of tasks accessing a resource through IO; the number of concurrent tasks cannot be limited. Therefore, in this scenario, `asyncio.Semaphore` is not a perfect solution. ### [asyncio.Queue](https://docs.python.org/3/library/asyncio-queue.html?ref=dataleadsfuture.com#asyncio.Queue) is the way Using `asyncio.Queue`, we can start a fixed number of concurrent tasks when the program starts, and then pass the data to be processed through the queue to these tasks. This is the well-known producer-consumer pattern. At the same time, like the multiprocessing queue, `asyncio.Queue` also serves to pass messages between concurrent tasks. --- ## The Magical World of asyncio’s Queue Why can `asyncio.Queue` plays such a significant role? In fact, we also encounter similar problems in real life: The most typical example is large shopping supermarkets. In such supermarkets, there are always many customers. After each customer finishesshopping, they need to check out. Checking out takes some time, which can lead to congestion. A more straightforward solution is to hire more cashiers, preferably one for each customer, for instant checkouts. However, this is unrealistic because so many cashiers would mean colossal cost pressure (and resource consumption) for the boss. So, a brilliant person came up with a good solution: have customers line up in a queue, and then have a few cashiers check out customers in turn. The only cost is that customers need to wait a little longer. At the same time, if the queue is too long, the manager can choose to temporarily add a few more cashiers. In this way, the system can flexibly expand. By comparing the customer queue to data entering the queue and cashiers to concurrent tasks, we can see the benefits `asyncio.Queue` brings: - It is a good implementation of the [producer-consumer pattern](https://en.wikipedia.org/wiki/Producer%E2%80%93consumer%5Fproblem?ref=dataleadsfuture.com). - It can control the number of concurrent tasks. - Making resource consumption manageable, and the system can also be flexibly expanded. --- ## The Adventures of the Producer-Consumer Pattern in asyncio ### What is the producer-consumer pattern Imagine two types of tasks sharing a queue. Task A produces data and puts it into the queue, while Task B retrieves data from the queue for processing. This is the producer-consumer pattern, where Task A is the producer, and Task B is the consumer. In analogy with a supermarket, customers are producers, cashiers are consumers, and the customer queue represents the queue. ### Why use the producer-consumer pattern In high-concurrency programs, producers often generate data quickly, while consumers process data slowly. Thus, producers must wait for consumers to finish processing before continuing to produce data. Sometimes, consumers process data quickly, while producers generate data slowly. This leads to consumers waiting for producers to generate data before continuing to run. To balance between producers and consumers, a queue is needed to store the data produced by the producer. The queue acts as a buffer and decouples the producer and consumer. ![The Diagram of Producer-Consumer Pattern.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-53.png) The Diagram of Producer-Consumer Pattern. Image by Author --- ## Implementing the Producer-Consumer Pattern with asyncio’s Queue Now, let’s implement the supermarket shopping scenario mentioned earlier using `asyncio.Queue`. ```Python import asyncio from asyncio import Queue from random import randrange class Product: def __init__(self, product_name: str, checkout_time: float): self.product_name = product_name self.checkout_time = checkout_time class Customer: def __init__(self, customer_id: int, products: list[Product]): self.customer_id = customer_id self.products = products ``` As shown in the code, we first implement the `Customer` and `Product` classes, representing customers and products that need to be checked out. The `Product` class has a `checkout_time` attribute, which represents the time required for checking out the product. After that, we implement a `checkout_customer` method that acts as a consumer. ```Python async def checkout_customer(queue: Queue, cashier_number: int): while not queue.empty(): customer: Customer = await queue.get() print(f"The Cashier_{cashier_number} " f"will checkout Customer_{customer.customer_id}") for product in customer.products: print(f"The Cashier_{cashier_number} " f"will checkout Customer_{customer.customer_id}'s " f"Product_{product.product_name}") await asyncio.sleep(product.checkout_time) print(f"The Cashier_{cashier_number} " f"finished checkout Customer_{customer.customer_id}") queue.task_done() ``` As long as there is data in the queue, this method will continue to loop. During each iteration, it uses a `get` method to retrieve a `Customer` instance. If there is no data in the queue, it will wait. After retrieving a piece of data (in this case, a `Customer` instance), it iterates through the `products` attribute and uses `asyncio.sleep` to simulate the checkout process. After finishing processing the data, we use `queue.task_done()` to tell the queue that the data has been successfully processed. Next, we implement the `generate_customer` method as a factory method for producing customers. We first define a product series and the required checkout time for each product. Then, we place 0 to 10 products in each customer’s shopping cart. ```Python def generate_customer(customer_id: int) -> Customer: all_products = [Product('deer', 2), Product('banana', .5), Product('sausage', .2), Product('diapers', .2)] products = [all_products[randrange(len(all_products))] for _ in range(randrange(10))] return Customer(customer_id, products) ``` Furthermore, we implement the `customer_generation` method as a producer. This method generates several customer instances regularly and puts them in the queue. If the queue is full, the `put` method will wait. ```Python async def customer_generation(queue: Queue): customer_count = 0 while True: customers = [generate_customer(the_id) for the_id in range(customer_count, customer_count+randrange(5))] for customer in customers: print("Waiting to put customer in line....") await queue.put(customer) print("Customer put in line...") customer_count = customer_count + len(customers) await asyncio.sleep(.3) ``` Finally, we use the `main` method to initialize the queue, producer, and consumer, and start all concurrent tasks. ```Python async def main(): customer_queue = Queue(2) customer_producer = asyncio.create_task(customer_generation(customer_queue)) cashiers = [checkout_customer(customer_queue, i) for i in range(3)] await asyncio.gather(customer_producer, *cashiers) if __name__ == "__main__": asyncio.run(main()) ``` ![The implementation is successful.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_UBOOK9UEqwqnCW1ugY6UxA.gif) The implementation is successful. Image by Author As expected, the implementation is successful. --- ## Introducing the PriorityQueue ### Why use [asyncio.PriorityQueue](https://docs.python.org/3/library/asyncio-queue.html?ref=dataleadsfuture.com#priority-queue) The queue mentioned earlier is a [First-In-First-Out (FIFO) queue](https://en.wikipedia.org/wiki/FIFO%5F%28computing%5Fand%5Felectronics%29?ref=dataleadsfuture.com), where the first item to enter the queue is the first to be retrieved. This is suitable when all tasks in the queue have the same priority. However, consider the following situation: Suppose there is a queue with tasks waiting in line, each requiring a long processing time. An error log or VIP user access is a high-priority task that needs immediate attention. What should we do? This is where [asyncio.PriorityQueue](https://docs.python.org/3/library/asyncio-queue.html?ref=dataleadsfuture.com#priority-queue) comes into play. ### Briefly describe asyncio.PriorityQueue’s implementation Unlike FIFO queues based on lists, `asyncio.PriorityQueue` is based on heaps. It is built using a binary tree structure. You may be familiar with binary search trees, which ensure that the most minor node is always the leftmost node. However, the binary tree in `asyncio.PriorityQueue` ensures that the most minor node is always at the top, so the highest priority node is permanently removed first. ![On the left is the binary tree used by PriorityQueue, and on the right is the binary search tree.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-56.png) On the left is the binary tree used by PriorityQueue, and on the right is the binary search tree. Image by Author ### Real-world example with asyncio.PriorityQueue Let’s illustrate the usage of `asyncio.PriorityQueue` with a real-world scenario that exists in practice. Imagine we have an order service API. The API takes time for each order to process, but we can’t keep users waiting too long. So when a user places an order, the API first puts the order into a queue, allowing a background task to process it asynchronously while immediately returning a message to the user. This API accepts orders from two types of users: regular users and VIP users. It must ensure that VIP user orders are processed with the highest priority. ![VIP orders are processed with the highest priority.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-57.png) VIP orders are processed with the highest priority. Image by Author To keep the learning curve low for readers, in this example, we will use `aiohttp` to implement the server. The specific code is as follows: ```Python import asyncio from asyncio import PriorityQueue, Task from dataclasses import dataclass, field from enum import IntEnum from random import randrange from aiohttp import web from aiohttp.web_app import Application from aiohttp.web_request import Request from aiohttp.web_response import Response app = Application() routers = web.RouteTableDef() QUEUE_KEY = "QUEUE_KEY" TASK_KEY = "TASK_KEY" class UserType(IntEnum): POWER_USER = 1 NORMAL_USER = 2 @dataclass(order=True) class WorkItem: user_type: UserType order_delay: int = field(compare=False) ``` First, we define an enumeration marking the two categories: regular users and VIP users. Next, we use `dataclass` to define a user's order, which contains the user type and order processing duration. The order duration is not considered in priority sorting. Then we define the consumer method `process_order_worker`, which retrieves orders from the queue and simulates the order processing. Don’t forget to use `queue.task_done()` to tell the queue that we finished processing the order. ```Python async def process_order_worker(worker_id: int, queue: PriorityQueue): while True: work_item: WorkItem = await queue.get() print(f"process_order_worker: Worker_{worker_id} begin to process worker {work_item}") await asyncio.sleep(work_item.order_delay) print(f"process_order_worker: Worker_{worker_id} finished to process worker {work_item}") queue.task_done() ``` Following that, we implement the order API using `aiohttp`. This API responds to user requests, generates an order object, and places it in the `asyncio.PriorityQueue`. It then immediately returns a response to the user, avoiding user wait time. ```Python @routers.post("/order") async def order(request: Request) -> Response: queue: PriorityQueue = app[QUEUE_KEY] body = await request.json() user_type = UserType.POWER_USER if body['power_user'] == 'True' else UserType.NORMAL_USER work_item = WorkItem(user_type, randrange(5)) await queue.put(work_item) return Response(body="order placed!") ``` When the program starts, we use `create_order_queue` to initialize the queue and order consumption tasks. ```Python async def create_order_queue(app: Application): print("create_order_queue: Begin to initialize queue and tasks.") queue: PriorityQueue = PriorityQueue(10) tasks = [asyncio.create_task(process_order_worker(i, queue)) for i in range(3)] app[QUEUE_KEY] = queue app[TASK_KEY] = tasks print("create_order_queue: Initialize queue and tasks success..") ``` When the program ends, we use `destroy_order_queue` to ensure that all orders in the queue are processed and the background tasks are closed correctly. `queue.join()` will wait for all the data in the queue to be processed. `asyncio.wait_for` sets a timeout of 20 seconds, after which it will no longer wait `queue.join()` to complete. ```Python async def destroy_order_queue(app: Application): queue: PriorityQueue = app[QUEUE_KEY] tasks: list[Task] = app[TASK_KEY] try: print("destroy_order_queue: Wait for 20 sec to let all work done.") await asyncio.wait_for(queue.join(), timeout=20.0) except Exception as e: print("destroy_order_queue: Cancel all tasks.") [task.cancel() for task in tasks] app.add_routes(routers) app.on_startup.append(create_order_queue) app.on_shutdown.append(destroy_order_queue) web.run_app(app) ``` We can test this implementation using PyCharm’s HTTP Request: ```Bash POST http://localhost:8080/order Content-Type: application/json {"power_user": "True"} ### POST http://localhost:8080/order Content-Type: application/json {"power_user": "False"} ### POST http://localhost:8080/order Content-Type: application/json {"power_user": "False"} ### POST http://localhost:8080/order Content-Type: application/json {"power_user": "True"} ``` ![API prioritizes orders from VIP users whenever possible.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-58.png) API prioritizes orders from VIP users whenever possible. Image by Author As you can see, the two high-priority tasks are processed as expected. Perfect! --- ## Conclusion In this article, I introduced the usage and best practices of `asyncio.Queue`. When you need to control the concurrency of a program, I recommend using `asyncio.Queue` to effectively manage resource consumption. I introduced the producer-consumer pattern and its benefits: 1. Balancing between producers and consumers, maximizing resource utilization. 2. Decoupling the system, allows producers and consumers to scale independently. Finally, I showed how to use `asyncio.PriorityQueue` to handle scenarios where tasks require prioritization through a real-world example. Due to space constraints, I could not cover all aspects of `asyncio.Queue`. However, I hope this article provided a solid understanding of the basic concepts and helpful examples. Asynchronous programming in Python is a powerful tool, and the producer-consumer pattern with `asyncio.Queue` is a versatile approach to handling concurrency and prioritization in your applications. ### Aiomultiprocess: Super Easy Integrate Multiprocessing & Asyncio in Python URL: https://www.dataleadsfuture.com/aiomultiprocess-super-easy-integrate-multiprocessing-asyncio-in-python/ Last updated: 2025-12-19T07:15:29.000Z In this article, I will introduce how to integrate [multiprocessing](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com) and [asyncio](https://docs.python.org/3/library/asyncio.html?ref=dataleadsfuture.com) using the [aiomultiprocess](https://aiomultiprocess.omnilib.dev/en/stable/?ref=dataleadsfuture.com) library easily. The article includes a [web scraping project](https://en.wikipedia.org/wiki/Web%5Fscraping?ref=dataleadsfuture.com) example and the best practices for using this library. --- ## Introduction My colleague Wang came to me today and said that his boss assigned him a new task: to write web scraping code to fetch information from 1,000 books on [books.toscrape.com](http://books.toscrape.com/?ref=dataleadsfuture.com) as quickly as possible. Wang told me: “I’ve read your related articles, and since the boss has performance requirements, why don’t I write one using [asyncio](https://docs.python.org/3/library/asyncio.html?ref=dataleadsfuture.com)? It doesn’t seem too difficult.” “30.09 seconds,” I told him a number. “What’s that?” Wang asked. I said I had just tried it and that only using concurrent tasks with asyncio for web scraping would take that long on my computer. This speed is already relatively fast. “12.64 seconds,” I told him another number. The speed doubled! Wang was stunned. Because I used a mighty library called `aiomultiprocess`, which can easily integrate multiprocessing and asyncio. And the performance can be improved by modifying the web scraping code with `aiomultiprocess` on the same network and computer. --- ## Multiprocessing and Asyncio: A Quick Recap Wang said: “That’s amazing! Teach me how to use aiomultiprocess quickly!” I told him not to hurry. Although the library is simple enough that he doesn’t need to understand what asyncio and multiprocessing are, I still need to give some theoretical introductions to help him truly master the implementation principles of this library. ### Key concepts of asyncio Asyncio is a new feature introduced in Python 3.4\. Its main part is to execute code snippets in a loop through an [event loop](https://docs.python.org/3/library/asyncio-eventloop.html?ref=dataleadsfuture.com) in the main thread. Users can switch to another task while waiting for a network call (or disk read/write) to return. Since it is single-threaded and not constrained by the [GIL](https://en.wikipedia.org/wiki/Global%5Finterpreter%5Flock?ref=dataleadsfuture.com), asyncio is very suitable for executing [IO-bound](https://en.wikipedia.org/wiki/I/O%5Fbound?ref=dataleadsfuture.com) code. ### Key concepts of multiprocessing [Multiprocessing](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com) is a feature introduced in Python for compute-intensive tasks. Its principle is to use multiple processes to execute different Python codes in parallel. It makes full use of the performance of multi-core CPUs, so it is very suitable for running [CPU-bound ](https://en.wikipedia.org/wiki/CPU-bound?ref=dataleadsfuture.com)code. ### Benefits of combining both approaches However, using asyncio or multiprocessing alone is only ideal in specific situations. In reality, the boundaries between IO-bound and CPU-bound tasks are not so clear. Take the web scraping scenario as an example: Web scraping is divided into two parts: fetching the HTML of the page from the network and parsing the required content from the HTML. The former is an IO-bound task, and the latter is a CPU-bound task. ![Web scraping contains both IO-bound and CPU-bound tasks.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-46.png) Web scraping contains both IO-bound and CPU-bound tasks. Image by Author This brings us a problem: if we only use asyncio, and the CPU time slice occupied by the CPU-bound task is too long, the single-threaded calculation performance will not be ideal. If we only use multiprocessing, the number of CPU cores will limit the concurrency. [In a previous article](https://medium.com/towards-data-science/harnessing-multi-core-power-with-asyncio-in-python-1764404ce44f?ref=dataleadsfuture.com), I introduced a way to integrate asyncio and multiprocessing. Specifically: In the main process’s event loop, use `loop.run_in_executor` to start multiple subprocesses. Then, use `asyncio.run` in each subprocess to create an event loop individually. The diagram is as follows: ![This diagram shows the way to integrate asyncio and multiprocessing.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-47.png) This diagram shows the way to integrate asyncio and multiprocessing. Image by Author However, this method has several problems: 1. It requires understanding more low-level asyncio and multiprocessing APIs, which is impractical. 2. It requires calculating which task to execute in which process initially without a flexible task allocation mechanism. 3. Due to mixing concurrent and parallel mechanisms, using queues or other methods to communicate between multiple tasks is challenging. 4. It is difficult to add new tasks during code execution. In summary, this method is too low-level and thus difficult to use. The `aiomultiprocess` library, which perfectly encapsulates the underlying code and exposes only a few upper-layer interfaces, can help us solve these problems well. --- ## Getting Started with Aiomultiprocess ### Installation and setup If you’re using pip: ```Bash Python -m pip install aiomultiprocess ``` If you’re using Anaconda since the default channel does not contain this package, use: ```Bash conda install -c conda-forge aiomultiprocess ``` ### Basic syntax and setup By examining the source code or referring to the [official documentation](https://aiomultiprocess.omnilib.dev/en/stable/?ref=dataleadsfuture.com), we can find that `aiomultiprocess` only requires three classes: 1. `Process`: Executes a coroutine task within a subprocess. It’s not commonly used, but the `Worker` and `Pool` classes inherit it. 2. `Worker`: Executes a coroutine task within a subprocess and returns the result. You can use this class to modify an existing coroutine function to run in a subprocess. 3. `Pool`: The core class we’ll be using. It aims to launch a process pool and allocate each coroutine task to a subprocess for execution. This class has two methods to master: - `map`: Takes a task function (coroutine function) and an iterable object as arguments. It applies each item in the iterable as an argument to the task function, running it in a subprocess. The method returns a generator object, and you can use `async for` to retrieve each value in the result. ```Python import asyncio import random import aiomultiprocess async def coro_func(value: int) -> int: await asyncio.sleep(random.randint(1, 3)) return value * 2 async def main(): results = [] async with aiomultiprocess.Pool() as pool: async for result in pool.map(coro_func, [1, 2, 3]): results.append(result) # The result depends on the order in which the parameters are passed in, # not on which task end first # Output: [2, 4, 6] print(results) if __name__ == "__main__": asyncio.run(main()) ``` - `apply`: Takes a task function, as well as `args` and `kwargs`. It combines the task function with `args` and `kwargs`, runs them in a subprocess, and returns an asyncio task. You can obtain all task results using `asyncio.gather`. ```Python import asyncio import random import aiomultiprocess async def coro_func(value: int) -> int: await asyncio.sleep(random.randint(1, 3)) return value * 2 async def main(): tasks = [] async with aiomultiprocess.Pool() as pool: tasks.append(pool.apply(coro_func, (1,))) tasks.append(pool.apply(coro_func, (2,))) tasks.append(pool.apply(coro_func, (3,))) results = await asyncio.gather(*tasks) print(results) # Output: [2, 4, 6] if __name__ == "__main__": asyncio.run(main()) ``` ### Understanding the key components Before diving into `aiomultiprocess` examples, we need to understand the implementation principles of the `Pool` class. `Pool` mainly consists of three modules: `scheduler`, `queue`, and `process`. Among them: 1. The `scheduler` module is responsible for task allocation. The default `scheduler` evenly distributes tasks across subprocesses in the order they are received. You can also implement a priority-based scheduler using `PriorityQueue`. 2. The `queue` module contains task and result queues, connecting the `scheduler` and `subprocesses`. Both queues are implemented using `multiprocessing.Manager().Queue()`. It is responsible for passing tasks to the subprocesses and returning results from the subprocesses. 3. The `process` module is implemented with the `Process` class, which acquires subprocesses through the [spawn](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#contexts-and-start-methods) method by default. An event loop is created in each subprocess. The results of the loop execution are passed back to the main process through the result queue. The entire schematic diagram is as follows: ![Aiomultiprocess consists of three components: scheduler, queue, and process.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-48.png) Aiomultiprocess consists of three components: scheduler, queue, and process. Image by Author --- ## Real-world Example: Web Scraping with Aiomultiprocess After introducing the basic usage and implementation principles of `aiomultiprocess`, I’ll now fulfill my promise to demonstrate how aiomultiprocess can easily improve existing code and achieve significant performance improvements. First, let me show you the asyncio version of the web scraping code. The goal of this code is simple: to fetch links to detail pages from 50 list pages on [books.toscrape.com](http://books.toscrape.com/?ref=dataleadsfuture.com), extract the desired book data from the detail pages, and finally write it to a CSV file: ```Python import asyncio import csv import time from string import Template from urllib.parse import urljoin from aiohttp import request from aiomultiprocess import Pool from bs4 import BeautifulSoup list_url_t = Template("https://books.toscrape.com/catalogue/category/books_1/page-$page.html") def get_detail_url(base_url: str, html: str) -> list[str]: """ Grab the link to the detail page of each book from the HTML code of the list page """ result = [] soup = BeautifulSoup(html, "html.parser") a_tags = soup.select("article.product_pod div.image_container a") for a_tag in a_tags: result.append(urljoin(base_url, a_tag.get("href"))) return result def parse_detail_page(html): """ Parse the HTML of the detail page to get the desired book data """ soup = BeautifulSoup(html, "lxml") title = soup.select_one("div.product_main h1").text price = soup.select_one("div.product_main p.price_color").text description_tag = soup.select_one("div#product_description + p") description = description_tag.text if description_tag else "" return {"title": title, "price": price, "description": description} async def fetch_list(url: str) -> list[str]: """ Get the URL of each detail page from the list page URL """ print(f"fetch_list: begin to process url: {url}") async with request("GET", url) as response: html = await response.text() urls = get_detail_url(url, html) return urls async def fetch_detail(url: str) -> dict: """ Get the book data on the detail page from the detail page URL """ async with request("GET", url) as response: html = await response.text() detail = parse_detail_page(html) return detail def write_to_csv(all_books: list): """ Writing data to CSV files """ print(f"write_to_csv: begin to write books detail to csv.") with open("../../raw_data/scraping_result.csv", "w", newline="", encoding="utf-8") as csv_file: fieldnames = all_books[0].keys() writer = csv.DictWriter(csv_file, fieldnames=fieldnames) writer.writerows(all_books) async def asyncio_main(): """ Implementing web scraping by using asyncio alone """ start = time.monotonic() all_books, detail_urls = [], [] fetch_list_tasks = [asyncio.create_task(fetch_list(list_url_t.substitute(page=i + 1))) for i in range(50)] for urls in asyncio.as_completed(fetch_list_tasks): detail_urls.extend(await urls) fetch_detail_tasks = [asyncio.create_task(fetch_detail(detail_url)) for detail_url in detail_urls] for detail in asyncio.as_completed(fetch_detail_tasks): all_books.append(await detail) write_to_csv(all_books) print(f"All done in {time.monotonic() - start} seconds") if __name__ == "__main__": asyncio.run(asyncio_main()) ``` ![The result of executing the asyncio version of the code.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-49.png) The result of executing the asyncio version of the code. Screenshot by Author Next, we only need to modify the `main` function, using aiomultiprocess’s `Pool` and corresponding APIs to start the web scraping code. The original logic code does not need to be changed: ```Python async def aiomultiprocess_main(): """ Integrating multiprocessing and asyncio with the help of aiomultiprocess, requires only a simple rewriting of the main function """ start = time.monotonic() all_books = [] async with Pool() as pool: detail_urls = [] async for urls in pool.map(fetch_list, [list_url_t.substitute(page=i + 1) for i in range(50)]): detail_urls.extend(urls) async for detail in pool.map(fetch_detail, detail_urls): all_books.append(detail) write_to_csv(all_books) print(f"All done in {time.monotonic() - start} seconds") if __name__ == "__main__": asyncio.run(aiomultiprocess_main()) ``` ![The result of the code execution of the aiomultiprocess version.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-50.png) The result of the code execution of the aiomultiprocess version. Screenshot by Author This way, a purely asyncio-based code is transformed into a version that integrates multiprocessing and asyncio. Isn’t it super simple? --- ## Best Practices for Using Aiomultiprocess `aiomultiprocess.Pool` provides parameters like `process`, `queuecount`, and `childconcurrency` for performance tuning. In the following sections, I’ll explain various methods of optimization based on different scenarios. ### Computation-intensive scenarios In this case, a task consumes significant processing time, leading to delayed IO responses when running high concurrency on a single process. To optimize, we can reduce the values of `childconcurrency` and `queuecount` while moderately increasing the `process`. ### High IO, high latency scenarios In this scenario, the IO pressure is relatively high, and the process needs to frequently switch between multiple IOs to promptly respond to IO returns. Here, we can increase `childconcurrency` to enhance each process’s concurrency. ### High concurrency, high throughput scenarios The `queue` is filled with numerous tasks in this situation, and multiple processes compete for `queue` resources. We can increase `queuecount `(but stay within the number of `processes`) to ensure that each process has sufficient queues for task allocation. ### Uneven task execution time scenarios Sometimes, some tasks require a lot of time for computation, such as parsing complex web pages. Others need substantial time to wait for IO, like writing massive amounts of data to a database. This results in task skew, with uneven pressure on processes executing tasks. Some processes quickly complete tasks and enter the waiting state, but others still need considerable time to finish the remaining tasks. In this case, we can execute different types of tasks separately, completing simple tasks first and tackling the more complex ones. Alternatively, we can lower the `queuecount` value to ensure fewer tasks enter the waiting queue, distributing tasks evenly among processes. We can also implement a custom `scheduler` to prioritize specific tasks, reducing task skew. --- ## Conclusion It’s best to combine concurrent and parallel code to maximize code performance. However, integrating `asyncio` and `multiprocessing` code requires interacting with many low-level APIs and writing a lot of foundational code, which can be daunting for many readers. `aiomultiprocess` solves this issue by enabling existing asyncio code to run on multiprocessing without any modifications. It simply utilizes its API calls, resulting in significant performance improvements. I hope you enjoy coding! If you’re interested in any of the points discussed in this article, please feel free to leave a comment and engage in discussion. ### Mastering Synchronization Primitives in Python Asyncio: A Comprehensive Guide URL: https://www.dataleadsfuture.com/mastering-synchronization-primitives-in-python-asyncio-a-comprehensive-guide/ Last updated: 2025-12-19T07:15:04.000Z In this article, I will introduce why you need synchronization primitives in Python’s asyncio and best practices for several synchronization primitives. And in the last part of the article, I will walk you through an example of synchronization primitives in action. --- ## Introduction ### Why do you need synchronization primitives in asyncio Anyone who has used Python multithreading knows that multiple threads share the same memory block. So when multiple threads perform non-atomic operations on the same area simultaneously, a thread-safe problem occurs. Since asyncio runs on a single thread, does it not have similar thread-safe issues? The answer is no. Concurrent tasks in asyncio are executed asynchronously, which means that there may be alternating execution of multiple tasks in time. A concurrency bug is triggered when one task accesses a particular memory area and waits for an IO operation to return, another task is also accessing this memory simultaneously. To avoid such bugs, Python asyncio introduces a synchronization primitive feature similar to multithreading. Also, to avoid too many tasks accessing a resource concurrently, asyncio’s synchronization primitives provide the ability to protect the resource by limiting the number of tasks accessing it simultaneously. Next, let’s take a look at what synchronization primitives are available in asyncio. --- ## Python Asyncio’s Synchronization Primitives ### [Lock](https://docs.python.org/3/library/asyncio-sync.html?ref=dataleadsfuture.com#asyncio.Lock) Before we introduce this API, let’s look at a situation: Suppose we have a concurrent task that needs a copy of the website data. It will first check if it’s in the cache; if it is, it will fetch it from the cache, and if not, it will read it from the website. Since it takes some time to read the website data to return and update the cache, when multiple concurrent tasks are executed at the same time, they all assume that this data does not exist in the cache and launch remote requests at the same time, as shown in the following code: ```Python import asyncio import aiohttp cache = dict() async def request_remote(): print("Will request the website to get status.") async with aiohttp.ClientSession() as session: response = await session.get("https://www.example.com") return response.status async def get_value(key: str): if key not in cache: print(f"The value of key {key} is not in cache.") value = await request_remote() cache[key] = value else: print(f"The value of key {key} is already in cache.") value = cache[key] print(f"The value of {key} is {value}") return value async def main(): task_one = asyncio.create_task(get_value("status")) task_two = asyncio.create_task(get_value("status")) await asyncio.gather(task_one, task_two) if __name__ == "__main__": asyncio.run(main()) ``` ![Both tasks think there is no data in the cache, thus accessing the remote site.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-39.png) Both tasks think there is no data in the cache, thus accessing the remote site. Image by Author Which is not in line with our original design intent, so `asyncio.Lock` comes in handy. We can check if there is data in the cache when concurrent tasks need to get a lock first, and other tasks that do not get a lock will wait. Until the task that gets the lock finishes updating the cache and releases the lock, the other tasks can continue to execute. The entire flowchart is shown below: ![The asyncio.Lock's entire flowchart.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-40.png) The asyncio.Lock's entire flowchart. Image by Author Let’s see how to write the code: ```Python import asyncio from asyncio import Lock import aiohttp cache = dict() lock = Lock() async def request_remote(): # ... async def get_value(key: str): async with lock: if key not in cache: print(f"The value of key {key} is not in cache.") value = await request_remote() cache[key] = value else: print(f"The value of key {key} is already in cache.") value = cache[key] print(f"The value of {key} is {value}") return value async def main(): # ... if __name__ == "__main__": asyncio.run(main()) ``` ![Only the first task needs to update the cache.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-41.png) Only the first task needs to update the cache. Image by Author Problem solved, isn’t it simple? ### [Semaphore](https://docs.python.org/3/library/asyncio-sync.html?ref=dataleadsfuture.com#semaphore) Sometimes, we need to access a resource with limited concurrent requests. For example, a particular database only allows five connections to be opened simultaneously. Or depending on the type of subscription you have, a web API only supports a certain number of concurrent requests at the same time. In this case, you need to use `asyncio.Semaphore`. `asyncio.Semaphore` uses an internal counter that decrements by one each time a Semaphore lock is acquired until it reaches zero. ![Semaphore will limit the number of concurrent tasks.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-42.png) Semaphore will limit the number of concurrent tasks. Image by Author When the counter of `asyncio.Semaphore` is zero, other tasks that need the lock will wait. When calling the release method after the execution of other tasks, the counter will be increased by one. The waiting tasks can continue to execute. The code example is as follows: ```Python import asyncio from asyncio import Semaphore from aiohttp import ClientSession async def get_url(url: str, session: ClientSession, semaphore: Semaphore): print('Waiting to acquire semaphore...') async with semaphore: print('Semaphore acquired, requesting...') response = await session.get(url) print('Finishing requesting') return response.status async def main(): # Although we start 1000 tasks, only 10 tasks will be executed at the same time. semaphore: Semaphore = Semaphore(10) async with ClientSession() as session: tasks = [asyncio.create_task(get_url("https://www.example.com", session, semaphore)) for _ in range(1000)] await asyncio.gather(*tasks) if __name__ == "__main__": asyncio.run(main()) ``` In this way, we can limit the number of connections that can be accessed concurrently. ### [BoundedSemaphore](https://docs.python.org/3/library/asyncio-sync.html?ref=dataleadsfuture.com#boundedsemaphore) Sometimes, due to code limitations, we can’t use `async with` to manage the acquire and release of semaphore locks, so we might call `acquire` somewhere and `release` somewhere else. What happens if we accidentally call the `asyncio.Semaphorerelease` method multiple times? ```Python import asyncio from asyncio import Semaphore async def acquire(semaphore: Semaphore): print("acquire: Waiting to acquire...") async with semaphore: print("acquire: Acquired...") await asyncio.sleep(5) print("acquire: Release...") async def release(semaphore: Semaphore): print("release: Releasing as one off...") semaphore.release() print("release: Released as one off...") async def main(): semaphore = Semaphore(2) await asyncio.gather(asyncio.create_task(acquire(semaphore)), asyncio.create_task(acquire(semaphore)), asyncio.create_task(release(semaphore))) await asyncio.gather(asyncio.create_task(acquire(semaphore)), asyncio.create_task(acquire(semaphore)), asyncio.create_task(acquire(semaphore))) if __name__ == "__main__": asyncio.run(main()) ``` As the code shows, we are limited to running two tasks simultaneously, but because we called release more than once, we can run three tasks at the same time next time. To solve this problem, we can use `asyncio.BoundedSemaphore` . As we know from the source code, when calling the `release`, a `ValueError` is thrown if the counter value is greater than the value set during initialization: ```Python import asyncio from asyncio import BoundedSemaphore async def main(): semaphore = BoundedSemaphore(2) await semaphore.acquire() semaphore.release() semaphore.release() if __name__ == "__main__": asyncio.run(main()) ``` ![When we call the release method multiple times, a ValueError is thrown.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-43.png) When we call the release method multiple times, a ValueError is thrown. Image by Author Therefore, the problem is being solved. ### [Event](https://docs.python.org/3/library/asyncio-sync.html?ref=dataleadsfuture.com#event) `Event` maintains an internal boolean variable as a flag. `asyncio.Event` has three common methods: `wait`, `set`, and `clear`. When the task runs to `event.wait()`, the task is in wait. At this point, you can call `event.set()` to set the internal marker to True, and all the waiting tasks can continue to execute. When the task is finished, you need to call `event.clear()` method to reset the value of the marker to False, to restore the event to its initial state, and you can continue to use the event next time. Instead of the sample code, I will show you how to use `Event` to implement an event bus at the end of the article. ### [Condition](https://docs.python.org/3/library/asyncio-sync.html?ref=dataleadsfuture.com#condition) `asyncio.Condition` is similar to `asyncio.Lock` and `asyncio.Event` combined. First, we will use `async with` to ensure that the condition lock is acquired, and then we call `condition.wait()` to release the condition lock and make the task wait temporarily. When `condition.wait()` passes, we regain the condition lock to ensure that only one task executes simultaneously. While a task temporarily releases the lock and goes into wait by `condition.wait()`, another task can either `async with` to the condition lock and notify all waiting tasks to continue execution by the `condition.notify_all()` method. The flowchart is shown below: ![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-44.png) The workflow of asyncio.Condition. Image by Author We can demonstrate the effect of `asyncio.Condition` with a piece of code: ```Python import asyncio from asyncio import Condition async def do_work(condition: Condition): print("do_work: Acquiring condition lock...") async with condition: print("do_work: Acquired lock, release and waiting for notify...") await condition.wait() print("do_work: Condition notified, re-acquire and do work.") await asyncio.sleep(1) print("do_work: Finished work, release condition lock.") async def fire_event(condition: Condition): await asyncio.sleep(5) print("fire_event: Acquiring condition lock....") async with condition: print("fire_event: Acquired lock, notify all workers.") condition.notify_all() print("fire_event: Notify finished, release the work...") async def main(): condition = Condition() asyncio.create_task(fire_event(condition)) await asyncio.gather(do_work(condition), do_work(condition)) if __name__ == "__main__": asyncio.run(main()) ``` Sometimes, we need `asyncio.Condition` to wait for a specific event to occur before proceeding to the next step. We can call the `condition.wait_for()` method and pass a method as an argument. Each time `condition.notify_all` is called, `condition.wait_for` checks the result of the execution of the parameter method and ends the wait if it is True, or continues to wait if it is False. We can demonstrate the effect of `wait_for` with an example. In the following code, we will simulate a database connection. Before executing the SQL statement, the code will check if the database connection is initialized and execute the query if the connection initialization is completed, or wait until the connection is completed initializing: ```Python import asyncio from asyncio import Condition from enum import Enum class ConnectionState(Enum): WAIT_INIT = 0 INITIALING = 1 INITIALIZED = 2 class Connection: def __init__(self): self._state = ConnectionState.WAIT_INIT self._condition = Condition() async def initialize(self): print("initialize: Preparing initialize the connection.") await self._change_state(ConnectionState.INITIALING) await asyncio.sleep(5) print("initialize: Connection initialized") await self._change_state(ConnectionState.INITIALIZED) async def execute(self, query: str): async with self._condition: print("execute: Waiting for connection initialized") await self._condition.wait_for(self._is_initialized) print(f"execute: Connection initialized, executing query: {query}") await asyncio.sleep(5) print("execute: Execute finished.") async def _change_state(self, state: ConnectionState): print(f"_change_state: Will change state from {self._state} to {state}") self._state = state print("_change_state: Change the state and notify all..") async with self._condition: self._condition.notify_all() def _is_initialized(self): if self._state is not ConnectionState.INITIALIZED: print("_is_initialized: The connection is not initialized.") return False print("_is_initialized: The connection is ready.") return True async def main(): connection = Connection() task_one = asyncio.create_task(connection.execute("SELECT * FROM table")) task_two = asyncio.create_task(connection.execute("SELECT * FROM other_table")) asyncio.create_task(connection.initialize()) await asyncio.gather(task_one, task_two) if __name__ == "__main__": asyncio.run(main()) ``` --- ## Some Tips for Using Synchronization Primitives ### Remember to use timeout or cancelation when needed When using synchronization primitives, we are generally waiting for the completion of a specific IO operation. However, due to network fluctuations or other unknown reasons, the IO operation of a task may take longer than others. In this case, we should set a timeout for the operation, so that when the execution time is too long, we can release the lock and allow other tasks to execute in time. In another case, we may loop through a task. It may keep some tasks waiting in the background and prevent the program from ending properly. At this point, remember to use cancel to terminate the cyclic execution of the task. ### Avoid using synchronization primitives or locking only the fewest resources We all know that the advantage of asyncio is that a task can switch to another task to execute while waiting for IO to return. But an asyncio task often contains both IO-bound operations and CPU-bound operations. If we lock too much code on the task, it will not be able to switch to another task in time, which will affect the performance. Therefore, if not necessary, try not to use synchronization primitives or only lock the least amount of resources. ### To avoid some other competitive locking situations There is no RLock in asyncio, so don’t use locks in recursive code. As with multithreading, asyncio also has the possibility of deadlocks, so try to avoid using multiple locks simultaneously. --- ## Advanced Techniques in Action: Asyncio-based Event Bus After the introduction earlier in the article, I believe you have a clear understanding of how to use asyncio’s synchronization primitives properly. Next, I will teach you how to use the synchronization primitives in real projects by taking you through the implementation of an event bus. As usual, the first step as an architect is to design the EventBus API. ```Python import asyncio from asyncio import Event import inspect from typing import Callable class EventBus: def __init__(self): self._event_dict = dict() async def on(self, event_name: str, fn: Callable): PASS def trigger(self, event_name: str, *args, **kwargs): PASS def _get_event(self, event_name: str): PASS ``` Since `EventBus` communicates using strings and internally, I intend to use `asyncio.Event` to implement the events corresponding to each string, we’ll start by implementing a `_get_event` method: ```Python def _get_event(self, event_name: str): if event_name in self._event_dict: print("event already inited...") event = self._event_dict.get(event_name) else: print(f"need to init a new event for {event_name}") event = Event() self._event_dict[event_name] = event return event ``` The `on` method will bind a callback function to a specific event: ```Python async def on(self, event_name: str, fn: Callable): event = self._get_event(event_name) while True: await event.wait() print("event fired") result = fn(*event.args, **event.kwargs) if inspect.isawaitable(result): await result # Since the callback function is likely a synchronous method, # we must perform an await here to allow other tasks to execute. await asyncio.sleep(0.1) event.clear() ``` The `trigger` method can manually trigger an event and pass in the corresponding data: ```Python def trigger(self, event_name: str, *args, **kwargs): event = self._get_event(event_name) event.args = args event.kwargs = kwargs event.set() ``` Finally, let’s write a `main` method to test the effect of EventBus: ```Python def a_sync_callback(data): print(f"A sync callback with data {data} is triggered") async def a_async_callback(data): await asyncio.sleep(1) print(f"A async callback with data {data} is triggered") async def main(): event_bus = EventBus() task_one = asyncio.create_task(event_bus.on("some_event", a_async_callback)) task_two = asyncio.create_task(event_bus.on("some_event", a_sync_callback)) event_bus.trigger("some_event", {id: 1}) await asyncio.wait([task_one, task_two], timeout=20.0) ``` At the end of the main method, remember to use timeout to prevent the program from executing all the time, as I warned before. ![The code is executed as expected.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-45.png) The code is executed as expected. Image by Author As you can see, the code is executed as expected. Isn’t it easy? --- ## Conclusion This article first introduced why Python asyncio needs synchronization primitives. Then, I introduced the best practices for Lock, Semaphore, Event, and Condition, and gave some tips on how to use them correctly. Finally, I have completed a small project with hands-on training on asyncio synchronization primitives, which I hope will help you better use synchronization primitives in real projects. Feel free to comment, share, or discuss topics about asyncio with me. ### Combining Traditional Thread-Based Code and Asyncio in Python URL: https://www.dataleadsfuture.com/combining-traditional-thread-based-code-and-asyncio-in-python/ Last updated: 2026-01-23T06:38:23.000Z In this article, I’ll explain how to call existing IO-blocking code in asyncio programs that don’t implement asyncio and how to call asyncio code in existing programs based on the threaded model. --- ## Introduction In the previous articles, I introduced you to asyncio, a Python feature. The performance of asyncio is very high, and using asyncio in modern, highly concurrent code will improve IO performance by several orders of magnitude. But in the real world, we have not seen asyncio code used as much as expected. Why is that? ### Challenge 1: How to call old IO-blocking code in asyncio code One scenario is that while we implement the new code with asyncio, there are still a lot of IO-blocking programs left in the system that was implemented traditionally. For example, microservice, file operation, etc. Even if you use asyncio and call these blocking APIs directly, you still can’t achieve the high concurrency effect. ### Challenge 2: How to call asyncio in existing blocking code to accomplish tasks asynchronously In another case, the existing code already implements a set of architecture based on the threaded model. Since asyncio’s event loop is executed in the current thread, directly calling asyncio will block the execution of existing code. It does not have the effect of concurrent execution. So today, I will use a few real-life examples to show you how to implement asyncio calls in each of the two cases. --- ## Part 1: Calling IO-Blocking Code in Asyncio-Based Programs Let’s take FastAPI as an example. FastAPI is a high-performance web framework based on asyncio implementation. But very often, not all the business logic of a web application is implemented in the FastAPI code. Sometimes we need to call several microservices implemented long ago that are blocking calls. How could we deal with this situation? ### Using run\_in\_executor to Run IO-Blocking Code In the previous article, we explained how to use `loop.run_in_executor` API to integrate multiple processes with asyncio to achieve high-performance computing. [Combining Multiprocessing and Asyncio in Python for Performance BoostsCombining Multiprocessing and asyncio via run\_in\_executor unifies the API for concurrent and parallel programming, simplifies our programming process, and allows us to obtain execution results in order of completion. This article will use a Real-world Example to Explain the Code Implementation![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_iQIahHdX4TP6KpbO.jpg)](https://www.dataleadsfuture.com/combining-multiprocessing-and-asyncio-in-python-for-performance-boosts/) However, IO-bound code is unsuitable for multi-process calls but recommended for multi-thread. The good thing is that the first argument of `loop.run_in_executor` accepts either a `concurrent.futures.ProcessPoolExecutor` implementation or a `concurrent.futures.ThreadPoolExecutor` implementation. So our sample code is as follows: ```Python import asyncio from concurrent.futures import ThreadPoolExecutor from typing import Callable, Any from fastapi import FastAPI import requests app = FastAPI() def get_status(q: str): """ Use the requests package to simulate microservice calls """ response = requests.get(f"https://www.example.com?q={q}") return response.status_code async def run_in_thread(fn: Callable, *args) -> Any: """ A tool method is used to call loop.run_in_executor """ task = app.state._loop.run_in_executor( app.state.executor, fn, *args ) return await task @app.get("/status") async def info(q: str): """ Calling synchronous methods within asynchronous methods """ status_code = await run_in_thread(get_status, (q,)) return f"The status is {status_code}" @app.on_event("startup") async def on_startup(): """ Initialize the thread pool and event loop at program startup """ app.state.executor = ThreadPoolExecutor() app.state._loop = asyncio.get_running_loop() @app.on_event("shutdown") async def on_shutdown(): """ Recycle resources when the program is shut down """ app.state.executor.shutdown() app.state._loop.close() ``` First, we use a `get_status` method to simulate the old microservice code calls via the requests package. Then, we manage the creation and destruction of the `ThreadPoolExecutor` thread pool during the startup and shutdown phases of the web application, respectively. Finally, we call the IO blocking methods in the thread pool and get the results in the response method of the request via `loop.run_in_executor`. The default executor parameter of `loop.run_in_executor` can be None. That is because asyncio will initialize a default thread pool internally after startup. When the executor parameter of `run_in_executor` is None, it will use the default thread pool to execute, so we don’t have to manage a thread pool in our code: ```Python async def run_in_thread(fn: Callable, *args) -> Any: """ A tool method is used to call loop.run_in_executor """ task = app.state._loop.run_in_executor( None, fn, *args ) return await task ``` ### Leveraging asyncio.to\_thread (Python 3.9+) Python 3.9 introduced a new high-level abstraction API, `asyncio.to_thread`, which, as you can see from the source code, internally calls the `loop.run_in_thread` method with the executor argument of None: ```Python @app.get("/status") async def info(q: str): """ Calling synchronous methods within asynchronous methods """ status_code = await asyncio.to_thread(get_status, (q,)) return f"The status is {status_code}" ``` Thus, using `asyncio.to_thread`, will further simplify the code. --- ### Part 2: Calling Asyncio Code in Traditional Thread-Based Programs There is another case where our program already implements a loop in the existing code. For example, most GUI programs use an event loop to respond to various events and to update the UI. Let’s take [tkinter](https://docs.python.org/3/library/tkinter.html?ref=dataleadsfuture.com) as an example. tkinter will start a main loop when it starts, and this main loop will block the main thread and keep on looping. As shown in the figure below: ![How does the tkinter main loop work.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-36.png) How does the tkinter main loop work. Image by Author ### A direct call to synchronous IO code will block the main loop Let’s take the example of a tkinter program that contains a button and a status text: ```Python import tkinter as tk import requests class App(tk.Tk): INIT_STATE = 0 # Initial state QUERYING_STATE = 1 # State when an HTTP request is started RESULT_STATE = 2 # The state of the HTTP request result def __init__(self): super().__init__() self.status_code = 0 self._refresh_ms = 60 self.state = App.INIT_STATE self._button = None self._label = None self.render_elements() self.after(self._refresh_ms, self.refresh) def render_elements(self): """ Set up the canvas and render UI elements """ self.geometry("400x150") self._button = tk.Button(self, text="request code", command=self.request_remote) self._label = tk.Label(self, text="") self._button.pack() self._label.pack() def request_remote(self): """ Initiate HTTP requests in a synchronous manner """ self.state = App.QUERYING_STATE response = requests.get("https://www.example.com") self.status_code = response.status_code self.state = App.RESULT_STATE def refresh(self): """ Refresh UI content every 60 milliseconds """ self.update_label() self.after(self._refresh_ms, self.refresh) def update_label(self): """ Update text content according to APP state """ match self.state: case App.INIT_STATE: self._label.config(text="Here will show the status code.") case App.QUERYING_STATE: self._label.config(text="Query remote...") case App.RESULT_STATE: self._label.config(text=f"The result code is: {self.status_code}") def start(self): self.mainloop() def main(): app = App() app.start() if __name__ == "__main__": main() ``` This program uses a state machine to implement it. Every 60 milliseconds, the code refreshes the corresponding text according to the program’s current state. When we click the request\_code button, the workflow should ideally look like the following diagram: ![Workflow of the tkinter program.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-37.png) Workflow of the tkinter program. Image by Author But by the execution result, the program hangs when clicking the button, and the status text is updated until the IO blocking code finishes executing. It means that the main loop is blocked when the IO request is running, causing the GUI interface to be unresponsive: ![The app is blocked and doesn’t show query text.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_yjV63hG98hbih_zp8D_GQA.gif) The app is blocked and doesn’t show query text. Image by Author ### Using asyncio.run to Run Asyncio Code Can we replace the requests package with the aiohttp package to achieve the asynchronous invocation of IO requests? Here we first inherit the `App` class to implement a new class `AppAsyncBase`. In this new class, we use aiohttp to implement an `async_request` method to lay the foundation for subsequent asynchronous calls: ```Python class AppAsyncBase(App): async def async_request(self): """ Asynchronous initiation of HTTP calls using aiohttp """ async with aiohttp.ClientSession() as session: async with session.get("https://www.example.com") as response: self.status_code = response.status self.state = App.RESULT_STATE ``` Readers of my previous article will know we can execute asynchronous methods inside synchronous code via `asyncio.run`: [Harnessing Multi-Core Power with Asyncio in PythonIn this article, I will show you how to execute Python asyncio code on a multi-core CPU to unlock the full performance of concurrent tasks.![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1__cicmvz8YXI0i6ouzF-mNA.webp)](https://www.dataleadsfuture.com/harnessing-multi-core-power-with-asyncio-in-python/) Then, we implement a new class `AppAsyncRun`, by inheriting `AppAsyncBase`. In this new class, we override the `request_remote` method and use `asyncio.run` to call the `async_request` method directly: ```Python class AppAsyncRun(AppAsyncBase): def __int__(self): super().__init__() def request_remote(self): """ Use asyncio.run to call the concurrent code """ self.state = AppAsyncBase.QUERYING_STATE asyncio.run(self.async_request()) ``` Next, let’s look at the results. Because asyncio’s event loop is executed in the main thread by default, and when the event loop is running, it blocks the main thread, and the main loop of tkinter is blocked and unresponsive: ![asyncio.run blocks the main loop.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_AvJkWEO1cyH17RmoEgGMBg.gif) asyncio.run blocks the main loop. Image by Author ### Integrating Asyncio with Thread-Based Programs Is there a way to solve the event loop blocking problem? Here we can use a separate daemon thread and then run the event loop into the daemon thread, so asyncio’s event loop will not block the main thread. The diagram is as follows: ![Combining tkinter and asyncio loops.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-38.png) Combining tkinter and asyncio loops. Image by Author Looking at the code implementation, we first inherit the `AppAsyncBase` class to implement a new class `AppEventLoop`. Next, override the `request_remote` method and use `asyncio.run_coroutine_threadsafe` to call the `async_request` method in the event loop. Request method in the event loop. `asyncio.run_coroutine_threadsafe` is also thread-safe: ```Python class AppEventLoop(AppAsyncBase): def __init__(self, loop: AbstractEventLoop = None): super().__init__() self._loop = loop def request_remote(self): """ Run the coroutine in event loop """ self.state = AppAsyncBase.QUERYING_STATE asyncio.run_coroutine_threadsafe(self.async_request(), self._loop) ``` Implement a `run_event_loop` method to call the `loop.run_forever` in the thread: ```Python def run_event_loop(loop: AbstractEventLoop): """ Run asyncio event loop in daemon thread """ loop.run_forever() ``` Then, use the `contextmanager` decorator to manage the lifecycle of the daemon thread: ```Python @contextmanager def get_event_loop(): """ Use a context manager to manage the thread's lifecycle""" loop = asyncio.get_event_loop() thread = Thread(target=run_event_loop, args=(loop,)) thread.daemon = True thread.start() try: yield loop finally: thread.join() ``` Finally, implement the event loop integration and the app launch in the main method, and let’s see the result: ```Python def main(): with get_event_loop() as loop: app = AppEventLoop(loop) app.start() ``` ![The event loop running separately in the daemon thread no longer blocks.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_xTGst08YrOmlVh3n2gFFKg.gif) The event loop running separately in the daemon thread no longer blocks. Image by Author Perfect! Click the button, the status text is changed accordingly, the whole GUI interface runs smoothly, and IO calls do not block the GUI ever. Mission accomplished. --- ## Conclusion Although asyncio can dramatically improve the execution performance of concurrent programs, asyncio is not used on a large scale because it does not implement much of the legacy code. Today’s article, using examples from real-world coding efforts, demonstrates the solution to two challenges: 1. How to call the old IO code non-blocking in a new asyncio program. 2. How to use asyncio asynchronous code in an existing synchronous program to achieve non-blocking execution. Welcome to leave your comments and discussions. I will answer them one by one. ### Harnessing Multi-Core Power with Asyncio in Python URL: https://www.dataleadsfuture.com/harnessing-multi-core-power-with-asyncio-in-python/ Last updated: 2025-03-26T02:07:56.000Z In this article, I will show you how to execute Python asyncio code on a multi-core CPU to unlock the full performance of concurrent tasks. --- ## Introduction ### What is our problem? asyncio uses only one core. In previous articles, I covered the mechanics of using Python asyncio in detail. With this knowledge, you can learn that asyncio allows IO-bound tasks to execute at high speed by manually switching task execution to bypass the GIL contention process during multi-threaded task switching. Theoretically, the execution time of IO-bound tasks depends on the time from initiation to the response of an IO operation and is not dependent on your CPU performance. Thus, we can concurrently initiate tens of thousands of IO tasks and complete them quickly. But recently, I was writing a program that needed to crawl tens of thousands of web pages simultaneously and found that although my asyncio program was much more efficient than programs that use iterative crawling of web pages, it still made me wait for a long time. Should I be using the full performance of my computer? So I opened Task Manager and checked: ![Only one core has a load.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-30.png) Only one core has a load. Image by Author I found that since the beginning, my code was running on only one CPU core, and several other cores were idle. In addition to launching IO operations to grab network data, a task has to unpack and format the data after it returns. Although this part of the operation does not consume much CPU performance, after more tasks, these CPU-bound operations will severely impact the overall performance. I wanted to make my asyncio concurrent tasks execute in parallel on multiple cores. Would that squeeze the performance out of my computer? --- ## The Underlying Principles of Asyncio To solve this puzzle, we must start with the underlying asyncio implementation, the event loop. ![How does the event loop work.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-31.png) How does the event loop work. Image by Author As shown in the figure, asyncio’s performance improvement for programs starts with IO-intensive tasks. IO-intensive tasks include HTTP requests, reading and writing files, accessing databases, etc. The most important feature of these tasks is that the CPU does not block and spends a lot of time computing while waiting for external data to be returned, which is very different from another class of synchronous tasks that require the CPU to be occupied all the time to compute a specific result. When we generate a batch of asyncio tasks, the code will first put these tasks into a queue. At this point, there is a thread called event loop that grabs one task at a time from the queue and executes it. When the task reaches the await statement and waits (usually for the return of a request), the event loop grabs another task from the queue and executes it. Until the previously waiting task gets data through a callback, the event loop returns to the previous waiting task and finishes executing the rest of the code. Since the event loop thread executes on only one core, the event loop blocks when the “rest of the code” happens to take up CPU time. When the number of tasks in this category is large, each small blocking segment adds up and slows down the program as a whole. --- ## What is My Solution From this, we know that asyncio programs slow down because our Python code executes the event loop on only one core, and the processing of IO data causes the program to slow down. Is there a way to start an event loop on each CPU core to execute it? As we all know, starting with Python 3.7, all asyncio code is recommended to be executed using the method `asyncio.run`, which is a high-level abstraction that calls the event loop to execute the code as an alternative to the following code: ```Python try: loop = asyncio.get_event_loop() loop.run_until_complete(task()) finally: loop.close() ``` As you can see from the code, each time we call `asyncio.run`, we get (if it already exists) or create a new event loop. Could we achieve our goal of executing asyncio tasks on multiple cores simultaneously if we could call the `asyncio.run` method on each core separately? The previous article used a real-life example to explain using asyncio’s `loop.run_in_executor` method to parallelize the execution of code in a process pool while also getting the results of each child process from the main process. If you haven’t read the previous article, you can check it out here: [Combining Multiprocessing and Asyncio in Python for Performance BoostsCombining Multiprocessing and asyncio via run\_in\_executor unifies the API for concurrent and parallel programming, simplifies our programming process, and allows us to obtain execution results in order of completion. This article will use a Real-world Example to Explain the Code Implementation![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_iQIahHdX4TP6KpbO.jpg)](https://www.dataleadsfuture.com/combining-multiprocessing-and-asyncio-in-python-for-performance-boosts/) Thus, our solution emerges: distribute many concurrent tasks to multiple sub-processes using multi-core execution via the `loop.run_in_executor` method, and then call `asyncio.run` on each sub-process to start the respective event loop and execute the concurrent code. The following diagram shows The entire flow: ![How the code is executed.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-32.png) How the code is executed. Image by Author Where the green part represents the sub-processes we started. The yellow part represents the concurrent tasks we started. --- ## Preparation Before Starting ### Simulating the task implementation Before we can solve the problem, we need to prepare before we start. In this example, we can’t write actual code to crawl the web content because it would be very annoying for the target website, so we will simulate our real task with code: ```Python async def fake_crawlers(): io_delay = round(random.uniform(0.2, 1.0), 2) await asyncio.sleep(io_delay) result = 0 for i in range(random.randint(100_000, 500_000)): result += i return result ``` As the code shows, we first use `asyncio.sleep` to simulate the return of the IO task in random time and an iterative summation to simulate the CPU processing after the data is returned. ### The effect of traditional code Next, we take the traditional approach of starting 10,000 concurrent tasks in a main method and watch the time consumed by this batch of concurrent tasks: ```Python async def main(): start = time.monotonic() tasks = [asyncio.create_task(fake_crawlers()) for i in range(10000)] await asyncio.gather(*tasks) print(f"All tasks completed. And last {time.monotonic() - start:.2f} seconds") ``` As the figure shows, executing the asyncio tasks with only one core takes a longer time. ![Takes a long time on a single core.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-33.png) Takes a long time on a single core. Image by Author --- ## The Code Implementation Next, let’s implement the multi-core asyncio code according to the flowchart and see if the performance is improved. ### Designing the overall structure of the code First, as an architect, we still need first to define the overall script structure, what methods are required, and what tasks each method needs to accomplish: ```Python import asyncio import time from concurrent.futures import ProcessPoolExecutor async def query_concurrently(begin_idx: int, end_idx: int): """ Start concurrent tasks by start and end sequence number """ def run_batch_tasks(batch_idx: int, step: int): """ Execute batch tasks in sub processes """ async def main(): """ Distribute tasks in batches to be executed in sub-processes """ ``` ### The specific implementation of each method Then, let’s implement each method step by step. The `query_concurrently` method will start the specified batch of tasks concurrently and get the results via the `asyncio.gather` method: ```Python sync def query_concurrently(begin_idx: int, end_idx: int): """ Start concurrent tasks by start and end sequence number """ tasks = [] for _ in range(begin_idx, end_idx, 1): tasks.append(asyncio.create_task(fake_crawlers())) results = await asyncio.gather(*tasks) return results ``` The `run_batch_tasks` method is not an async method, as it is started directly in the child process: ```Python def run_batch_tasks(batch_idx: int, step: int): """ Execute batch tasks in sub processes """ begin = batch_idx * step + 1 end = begin + step results = [result for result in asyncio.run(query_concurrently(begin, end))] return results ``` Finally, there is our `main` method. This method will call the `loop.run_in_executor` method to have the `run_batch_tasks` method execute in the process pool and merge the results of the child process execution into a list: ```Python async def main(): """ Distribute tasks in batches to be executed in sub-processes """ start = time.monotonic() loop = asyncio.get_running_loop() with ProcessPoolExecutor() as executor: tasks = [loop.run_in_executor(executor, run_batch_tasks, batch_idx, 2000) for batch_idx in range(5)] results = [result for sub_list in await asyncio.gather(*tasks) for result in sub_list] print(f"We get {len(results)} results. All last {time.monotonic() - start:.2f} second(s)") ``` Since we are writing a multi-process script, we need to use `if __name__ == “__main__”` to start the main method in the main process: ```Python if __name__ == "__main__": asyncio.run(main()) ``` ### Execute the code and see the results Next, we start the script and look at the load on each core in the task manager: ![All cores are almost utilized.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-34.png) All cores are almost utilized. Image by Author As you can see, all the CPU cores are utilized. Finally, we observe the code execution time and confirm that the multi-threaded asyncio code does indeed speed up the code execution by several times! Mission accomplished! ![Nearly triple the performance boost!](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-35.png) Nearly triple the performance boost! Image by Author --- ## Conclusion In this article, I explained why asyncio could execute IO-intensive tasks concurrently but still takes longer than expected when running large batches of concurrent tasks. It is because in the traditional implementation scheme of asyncio code, the event loop can only execute tasks on one core, and the other cores are in an idle state. So I have implemented a solution for you to call each event loop on multiple cores separately to execute concurrent tasks in parallel. And finally, it improved the performance of the code significantly. Due to the limitation of my ability, the solution in this article inevitably has imperfections. I welcome your comments and discussion. I will actively answer for you. ### Combining Multiprocessing and Asyncio in Python for Performance Boosts(Updated in 2026) URL: https://www.dataleadsfuture.com/combining-multiprocessing-and-asyncio-in-python-for-performance-boosts/ Last updated: 2026-07-08T00:57:27.000Z Hey, I'm honestly flattered that you're still reading this article I wrote three years ago. So I went back, took a fresh look at the whole thing, and updated some key parts. While I was at it, I gave the code a bit of a makeover, fixed a few small issues, and updated some APIs for newer Python versions to keep things compatible. Let's jump right in. ## Introduction Thanks to [GIL](https://wiki.python.org/moin/GlobalInterpreterLock?ref=dataleadsfuture.com), using multiple threads to perform CPU-bound tasks has never been an option. With the popularity of multicore CPUs, Python offers a [multiprocessing](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com) solution to perform CPU-bound tasks. But until now, there were still some problems with using multiprocess-related APIs directly. Before we start, we still have a small piece of code to aid in the demonstration: ```Python import time from multiprocessing import Process def sum_to_num(final_num: int) -> int: start = time.monotonic() result = 0 for i in range(0, final_num+1, 1): result += i print(f"The method with {final_num} completed in {time.monotonic() - start:.2f} second(s).") return result ``` The method takes one argument and starts accumulating from 0 to this argument. Print the method execution time and return the result. ### Problems with [multiprocessing](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#the-process-class) ```Python def main(): # We initialize the two processes with two parameters, from largest to smallest process_a = Process(target=sum_to_num, args=(200_000_000,)) process_b = Process(target=sum_to_num, args=(50_000_000,)) # And then let them start executing process_a.start() process_b.start() # Note that the join method is blocking and gets results sequentially start_a = time.monotonic() process_a.join() print(f"Process_a completed in {time.monotonic() - start_a:.2f} seconds") # Because when we wait process_a for join. The process_b has joined already. # so the time counter is 0 seconds. start_b = time.monotonic() process_b.join() print(f"Process_b completed in {time.monotonic() - start_b:.2f} seconds") ``` As the code shows, we directly create and start multiple processes, and call the start and join methods of each process. However, there are some problems here: 1. The `join` method cannot return the result of task execution. 2. The `join` method blocks the main process and executes it sequentially. Even if the later tasks are executed faster than the earlier ones, as shown in the following figure: ![The screenshot shows the execution sequence of join.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-21.png) The screenshot shows the execution sequence of join. Image by Author ![Although process_b finishes executing first, it still has to wait for process_a.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-22.png) Although process\_b finishes executing first, it still has to wait for process\_a. Image by Author ### Problems of using [Pool](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#module-multiprocessing.pool) If we use `multiprocessing.Pool`, there are also some problems: ```Python def main(): with Pool() as pool: result_a = pool.apply(sum_to_num, args=(200_000_000,)) result_b = pool.apply(sum_to_num, args=(50_000_000,)) print(f"sum_to_num with 200_000_000 got a result of {result_a}.") print(f"sum_to_num with 50_000_000 got a result of {result_b}.") ``` As the code shows, Pool’s [apply](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#multiprocessing.pool.Pool.apply) method is synchronous, which means you have to wait for the previously apply task to finish before the next `apply` task can start executing. ![multiprocessing.Pool.apply method is synchronous.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-23.png) [multiprocessing.Pool.apply](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#module-multiprocessing.pool) method is synchronous. Image by Author Of course, we can use the [apply\_async](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#multiprocessing.pool.Pool.apply%5Fasync) method to create the task asynchronously. But again, you need to use the get method to get the result blockingly. It brings us back to the problem with the join method: ```Python def main(): with Pool() as pool: result_a = pool.apply_async(sum_to_num, args=(200_000_000,)) result_b = pool.apply_async(sum_to_num, args=(50_000_000,)) print(f"sum_to_num with 200_000_000 got a result of {result_a.get()}.") print(f"sum_to_num with 50_000_000 got a result of {result_b.get()}.") ``` ![Although apply_async is asynchronous, get will still block and execute sequentially.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-24.png) Although [apply\_async](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#multiprocessing.pool.Pool.apply%5Fasync) is asynchronous, get will still block and execute sequentially. Image by Author ### The problem with using [ProcessPoolExecutor](https://docs.python.org/3/library/concurrent.futures.html?ref=dataleadsfuture.com#concurrent.futures.Executor) directly So, what if we use `concurrent.futures.ProcesssPoolExecutor` to execute our CPU-bound tasks? ```Python def main(): with ProcessPoolExecutor() as executor: numbers = [200_000_000, 50_000_000] for result in executor.map(sum_to_num, numbers): print(f"sum_to_num got a result which is {result}.") ``` As the code shows, everything looks great and is called just like `asyncio.as_completed`. But look at the results; they are still fetched in startup order. This is not at all the same as `asyncio.as_completed`, which gets the results in the order in which they were executed: ![Results are fetched in startup order.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-25.png) Results are fetched in startup order. Image by Author ![The result of the iteration still maintains the call order and blocks.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-26.png) The result of the iteration still maintains the call order and blocks. Image by Author ### Use asyncio’s [run\_in\_executor](https://docs.python.org/3/library/asyncio-eventloop.html?ref=dataleadsfuture.com#asyncio.loop.run%5Fin%5Fexecutor) to fix it Fortunately, we can use asyncio to handle IO-bound tasks, and its `run_in_executor` method to invoke multi-process tasks in the same way as asyncio. Not only unifying concurrent and parallel APIs, but also solving the various problems we encountered above: ```Python async def main(): loop = asyncio.get_running_loop() tasks = [] with ProcessPoolExecutor() as executor: for number in [200_000_000, 50_000_000]: tasks.append(loop.run_in_executor(executor, sum_to_num, number)) # Or we can just use the method asyncio.gather(*tasks) for done in asyncio.as_completed(tasks): result = await done print(f"sum_to_num got a result which is {result}") ``` ![Combining asyncio and ProcessPoolExecutor.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-27.png) Combining asyncio and ProcessPoolExecutor. Image by Author Since the sample code in the previous article was all about simulating what we should call the methods of the concurrent process, many readers still need help understanding how to use it in actual coding after learning it. So, after understanding why we need to perform CPU-bound parallel tasks in asyncio, today we will use a real-world example to explain how to use asyncio to handle IO-bound and CPU-bound tasks simultaneously and appreciate the efficiency of asyncio for our code. Note: Before continuing, if you are interested in the practice of using `asyncio.gather` and `asyncio.as_completed`, you can read this article of mine: [Use These Methods to Make Your Python Concurrent Tasks Perform BetterBest practices for asyncio.gather, asyncio.as\_completed, and asyncio.wait![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_Eel1JPfhnioG71TL.jpg)](https://www.dataleadsfuture.com/use-these-methods-to-make-your-python-concurrent-tasks-perform-better/) --- ### Real-world Case: Concurrent File Reading and MapReduce Data Processing In this case today, we will deal with two problems: 1. How to read multiple datasets concurrently. Especially if the datasets are large or numerous. How to use asyncio to improve efficiency. 2. How to use asyncio’s `run_in_executor` method to implement a MapReduce program and process datasets efficiently. Before we start, I will explain to you how our code is going to be executed using a diagram: ![The diagram shows how the entire code works.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-28.png) The diagram shows how the entire code works. Image by Author The yellow part represents our concurrent tasks. Since the CPU can process data from memory faster than IO can read data from disk, we first read all datasets into memory concurrently. After the initial data merging and slicing, we come to the green part that represents the CPU parallel task. In this part, we will start several processes to map the data. Finally, we get the intermediate results of all the processes in the main process and then use a `reduce` program to get the final results. Of course, if you're looking for an out-of-the-box solution, Aiomultiprocess is what you need. Here's a detailed article about it: [Supercharge Your Python Asyncio With Aiomultiprocess: A Comprehensive GuideHarness the power of asyncio and multiprocessing to turbocharge your applications![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_8QGPKl2DyKWYxsKVj_zzZw.webp)](https://www.dataleadsfuture.com/supercharge-your-python-asyncio-with-aiomultiprocess-a-comprehensive-guide/) --- ## Data Preparation and Installation of Dependencies ### Data preparation In this case, we will use the [Google Books Ngram Dataset](https://books.google.com/ngrams/info?ref=dataleadsfuture.com), which counts the frequency of each string combination in various books by year from 1500 to 2012. The Google Books Ngram dataset is free for any purpose, and today we will use the following datasets: - [http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-a.gz](http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-a.gz?ref=dataleadsfuture.com) - [http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-b.gz](http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-b.gz?ref=dataleadsfuture.com) - [http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-c.gz](http://storage.googleapis.com/books/ngrams/books/googlebooks-eng-all-1gram-20120701-c.gz?ref=dataleadsfuture.com) We aim to count the cumulative number of times each word is counted by the result set. ### Dependency installation To read the files concurrently, we will use the [aiofiles](https://pypi.org/project/aiofiles/?ref=dataleadsfuture.com) library, which can support asyncio’s concurrent implementation. If you are using pip, you can install it as follows: ```Bash $ pip install aiofiles ``` If you are using Anaconda, you can install it as follows: ```Bash $ conda install -c anaconda aiofiles ``` --- ## Code Structure Design Since this case is still relatively simple, for the sake of demonstration, we will use a `.py` script to do the whole thing here. As an architect, before you start, you should plan your methods according to the flowchart design and try to follow the “single responsibility principle” for each method. Thus, do only one thing once upon each method: ```Python import asyncio from concurrent.futures import ProcessPoolExecutor from itertools import batched from collections.abc import Iterator import functools import time from tqdm import tqdm from tqdm.asyncio import tqdm_asyncio from aiofiles import open async def read_file(filename: str): """ We will use the aiofiles API to read the files concurrently """ pass async def get_all_file_content(file_names: list[str]): """ Start concurrent tasks and join the file contents together """ pass def partition(contents: list[str], partition_size: int) -> Iterator[tuple[str, ...]]: """ Split the contents into multiple lists of partition_size length and return them as generator """ pass def map_resource(chunk: list[str]) -> dict[str, int]: """ The method that actually performs the map task returns the sum of the counts corresponding to the keywords in the current partition. """ pass def map_with_process(chunks: list[list[str]]): """ Execute map tasks in parallel and join the results of multiple processes into lists """ pass async def merge_resource(first: dict[str, int], second: dict[str, int]) -> dict[str, int]: """ The actually reduce method sums the counts of two dicts with the same key """ pass def reduce(intermediate_results: list[dict[str, int]]) -> dict[str, int]: """ Use the functools.reduce method to combine all the items in the list """ pass async def main(partition_size: int): """ Entrance to all methods """ pass if __name__ == "__main__": asyncio.run(main(partition_size=60_000)) ``` --- ## Code Implementation Next, we will implement each method step by step and finally integrate them to run together in the `main` method. ### File reading Method `read_file` will implement reading a single file with `aiofiles`: ```Python async def read_file(filename: str): """ We will use the aiofiles API to read the files concurrently """ async with open(filename, "r", encoding="utf-8") as f: return await f.readlines() ``` Method `get_all_file_content` will start the file reading task and, after all the files have been read, will merge each line of text into a list and return it. ```Python async def get_all_file_content(file_names: list[str]): """ Start concurrent tasks and join the file contents together """ print("Begin to read files...") start = time.monotonic() tasks = [] for filename in file_names: tasks.append(asyncio.create_task(read_file(filename))) temp_results = await asyncio.gather(*tasks) results = [item for sublist in temp_results for item in sublist] print(f"All files are read in {time.monotonic() - start:.2f} second(s)") return results ``` ### Data grouping Method `partition` will decompose the list into multiple smaller lists of partition\_size length according to the passed partition\_size and facilitate subsequent iterations using the generator: ```Python def partition(contents: list[str], partition_size: int) -> Iterator[tuple[str, ...]]: """ Split the contents into multiple lists of partition_size length and return them as generator """ yield from batched(contents, partition_size) ``` ### Map processing data Method `map_resource` is the actual map method. Use it to read each line of data from the list, use the word as the key and the sum of the frequencies as the value, and finally return a dict result. ```Python def map_resource(chunk: list[str]) -> dict[str, int]: """ The method that actually performs the map task returns the sum of the counts corresponding to the keywords in the current partition. """ result = {} for line in chunk: word, _, count, _ = line.split('\t') if word in result: result[word] = result[word] + int(count) else: result[word] = int(count) return result ``` ### Integrating asyncio with multiprocessing Method `map_with_process` calls asyncio’s `run_in_executor` method, which starts a pool of processes according to the number of CPU cores and executes the map method in parallel. And the final result is merged into a list by `asyncio.gather` method. ```Python async def map_with_process(chunks: list[list[str]]): """ Execute map tasks in parallel and join the results of multiple processes into lists """ print("Start parallel execution of map tasks...") start = time.monotonic() loop = asyncio.get_running_loop() tasks = [] with ProcessPoolExecutor() as executor: for chunk in chunks: tasks.append(loop.run_in_executor(executor, map_resource, chunk)) print(f"All map tasks are executed in {time.monotonic()-start:.2f} second(s)") return await asyncio.gather(*tasks) ``` ### Reducing the merged data Since the previous map process ends up with a list of word frequencies processed by multiple processes, we also need to use a `reduce` method to merge numerous dicts into a single final result, recording the final frequency of each word. Here we first write the method implementation of the `reduce` process. ```Python def merge_resource(first: dict[str, int], second: dict[str, int]) -> dict[str, int]: """ The actually reduce method sums the counts of two dicts with the same key """ merged = first for key in second: if key in merged: merged[key] = merged[key] + second[key] else: merged[key] = second[key] return merged ``` Then we call the `functools.reduce` method directly to merge the data. ```Python def reduce(intermediate_results: list[dict[str, int]]) -> dict[str, int]: """ Use the functools.reduce method to combine all the items in the list """ return functools.reduce(merge_resource, intermediate_results) ``` ### Finally, implement the main method Eventually, we will integrate all the methods into the `main` method and call. ```Python async def main(partition_size: int): """ Entrance to all methods """ file_names = ["../data/googlebooks-eng-all-1gram-20120701-a", "../data/googlebooks-eng-all-1gram-20120701-b", "../data/googlebooks-eng-all-1gram-20120701-c"] contents = await get_all_file_content(file_names) chunks = partition(contents, partition_size) intermediate_results = await map_with_process(chunks) final_results = reduce(intermediate_results) print(f'Aardvark has appeared {final_results["Aardvark"]} times.') ``` Great! We get the sum of the frequencies of the word Aardvark in all the datasets. Task complete. ### Using tqdm to indicate progress In the previous article, we explained how to use [tqdm](https://github.com/tqdm/tqdm?ref=dataleadsfuture.com) to indicate the progress of asyncio tasks. [Using Tqdm with Asyncio in PythonAn efficient way to monitor concurrent tasks’ progress![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_L303gsPJ5jU2Iokh.jpg)](https://www.dataleadsfuture.com/using-tqdm-with-asyncio-in-python/) Since in the real world, data processing of large datasets often takes a long time, during which we need to track the progress of code execution, we also need to add `tqdm` progress bars in the right places. ```Python async def get_all_file_content(file_names: list[str]): """ Start concurrent tasks and join the file contents together """ print("Begin to read files...") start = time.monotonic() tasks, results = [], [] for filename in file_names: tasks.append(asyncio.create_task(read_file(filename))) temp_results = await tqdm_asyncio.gather(*tasks) # add tqdm asyncio API results = [item for sublist in temp_results for item in sublist] print(f"All files are read in {time.monotonic() - start:.2f} second(s)") return results ######################################### async def map_with_process(chunks: list[list[str]]): """ Execute map tasks in parallel and join the results of multiple processes into lists """ print("Start parallel execution of map tasks...") start = time.monotonic() loop = asyncio.get_running_loop() tasks = [] with ProcessPoolExecutor() as executor: for chunk in chunks: tasks.append(loop.run_in_executor(executor, map_resource, chunk)) print(f"All map tasks are executed in {time.monotonic()-start:.2f} second(s)") return await tqdm_asyncio.gather(*tasks) # add tqdm asyncio API ########################################### def reduce(intermediate_results: list[dict[str, int]]) -> dict[str, int]: """ Use the functools.reduce method to combine all the items in the list """ return functools.reduce(merge_resource, tqdm(intermediate_results)) ``` It looks much more professional now. ![The resulting screenshot after adding the tqdm APIs.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-29.png) The resulting screenshot after adding the tqdm APIs. Image by Author --- ## Follow Data Leads Future One practical story every month, sharing my hard-learned experiences in the enterprise AI space. Subscribe Email sent! Check your inbox to complete your signup. You can unsubscribe anytime. ## Conclusion In today’s article, we explored some of the problems with multi-process code, such as the hassle of getting the results of each process and the inability to get the results in the order in which we execute the tasks. We also explored the feasibility of integrating asyncio with `ProcessPoolExecutor` and the advantages that such integration brings to us. For example, it unifies the API for concurrent and parallel programming, simplifies our programming process, and allows us to obtain execution results in order of completion. Finally, we explain how we can alternate between concurrent and parallel programming techniques to help us execute our code efficiently in data science tasks through a real-world case study that exists. Due to the limited ability of individuals, there are inevitably few places in this case, so I welcome your comments and corrections so that we can learn and progress together. --- next step Author's Recommendation: Optimizing asyncio and multiprocessing is useless if your underlying code still chokes on synchronous File I/O. To write true production-grade backend systems, I recommend the [****Introduction to Programming with Python Specialization**](https://imp.i384100.net/GbQoyn?ref=dataleadsfuture.com) on Coursera. This program focuses on file system traversal, robust exception handling, and low-level scripting. **If you choose to enroll, I may earn a small commission at zero extra cost to you. I only recommend high-quality resources that genuinely align with the engineering standards of Data Leads Future.* [Start Learning for Free Now ](https://imp.i384100.net/GbQoyn?ref=dataleadsfuture.com) ### Using Tqdm with Asyncio in Python URL: https://www.dataleadsfuture.com/using-tqdm-with-asyncio-in-python/ Last updated: 2025-12-19T07:13:13.000Z ## Introduction ### What’s bothering me Using concurrent programming in Python for efficiency is not unusual for a data scientist. Watching various sub-processes or concurrent threads in the background to keep my computation or IO-bound tasks in order is always satisfying. But one thing that still bothers me is that when I’m concurrently processing hundreds or thousands of files or executing hundreds of processes in the background, I’m always worried about whether a few tasks will hang secretly and the whole code will never finish. I also have difficulty knowing where the code is now in execution. The worst part is that when I’m looking at a blank screen, it’s hard to tell how much longer my code will take to execute or what the ETA is. This is very detrimental to my ability to organize my work schedule. Therefore, I wanted a way to let me know where the code execution had gotten to. ### How it was done in the past A more traditional approach is to [share a memory area](https://docs.python.org/3/library/multiprocessing.html?ref=dataleadsfuture.com#exchanging-objects-between-processes) between tasks, put a counter in this memory area, let this counter+1 when a task is finished, and then use a thread to keep printing the value of this counter. This is never a good solution: On the one hand, I need to add a code for counting into your existing business logic, which violates the principle of “low coupling, high cohesion”. On the other hand, I’d have to be very careful with the locking mechanism due to thread-safety issues, which would cause unnecessary performance problems. ### tqdm is the way ![tqdm uses a progress bar to indicate the progress of your tasks.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_khVTOy3hSOi1McJ6hvlhSg.gif) tqdm uses a progress bar to indicate the progress of your tasks. Image by Author One day, I discovered the [tqdm library](https://github.com/tqdm/tqdm?ref=dataleadsfuture.com), which uses a progress bar to visualize the progress of my code. Could I use the progress bar to visualize the completion and ETA of my asyncio tasks? I went ahead and researched, and I succeeded. Then I’m sharing this method with you so that every programmer can have a chance to monitor their concurrent tasks' progress. Let’s go. --- ## Background on asyncio in Python Before we start, I’d like you to get some background on Python asyncio. My article describes the usage of some of asyncio’s common APIs, which will help us better understand the design of tqdm: [Use These Methods to Make Your Python Concurrent Tasks Perform BetterBest practices for asyncio.gather, asyncio.as\_completed, and asyncio.wait![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_Eel1JPfhnioG71TL.jpg)](https://www.dataleadsfuture.com/use-these-methods-to-make-your-python-concurrent-tasks-perform-better/) --- ## Overview of tqdm As the [official website](https://tqdm.github.io/?ref=dataleadsfuture.com) describes, tqdm is a tool that displays a progress bar for your loops. It is straightforward to use, highly customizable and has a shallow resource footprint. A typical usage is to pass an iterable object into the tqdm constructor, and you get a progress bar like the following: ```Python from time import sleep from tqdm import tqdm def main(): for _ in tqdm(range(100)): # do something in the loop sleep(0.1) if __name__ == "__main__": main() ``` Or you can manually go through and update the progress of the progress bar as the file is being read: ```Python import os from tqdm import tqdm def main(): filename = "../data/large-dataset" with (tqdm(total=os.path.getsize(filename)) as bar, open(filename, "r", encoding="utf-8") as f): for line in f: bar.update(len(line)) if __name__ == "__main__": main() ``` ![Use tqdm to indicate the progress of reading a large dataset.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_5pyEGs7fnBsYhF1N3J9iBg.gif) Use tqdm to indicate the progress of reading a large dataset. Image by Author --- ## Integrating tqdm with asyncio Overall, tqdm is very easy to use. However, there needs to be more information on GitHub about integrating tqdm with asyncio. So I went digging through the source code to see if tqdm supports asyncio. Fortunately, the latest version of tqdm provides the package `tqdm.asyncio`, which provides the Class `tqdm_asyncio`. The Class `tqdm_asyncio` has two related methods. One is `tqdm_asyncio.as_completed`. As you can see from the source code, it is a wrapper for `asyncio.as_completed`: ```Python @classmethod def as_completed(cls, fs, *, loop=None, timeout=None, total=None, **tqdm_kwargs): """ Wrapper for `asyncio.as_completed`. """ if total is None: total = len(fs) kwargs = {} if version_info[:2] < (3, 10): kwargs['loop'] = loop yield from cls(asyncio.as_completed(fs, timeout=timeout, **kwargs), total=total, **tqdm_kwargs) ``` The other is `tqdm_asyncio.gather` , which, as seen from the source code, is based on an implementation of `tqdm_asyncio.as_completed` that emulates the functionality of `asyncio.gather`: ```Python @classmethod async def gather(cls, *fs, loop=None, timeout=None, total=None, **tqdm_kwargs): """ Wrapper for `asyncio.gather`. """ async def wrap_awaitable(i, f): return i, await f ifs = [wrap_awaitable(i, f) for i, f in enumerate(fs)] res = [await f for f in cls.as_completed(ifs, loop=loop, timeout=timeout, total=total, **tqdm_kwargs)] return [i for _, i in sorted(res)] ``` So, next, I will describe the usage of these two APIs. Before we start, we also need to do some preparation work. Here, I have written a simple method that simulates a concurrent task with a random sleep time: ```Python import asyncio import random from tqdm.asyncio import tqdm_asyncio class AsyncException(Exception): def __int__(self, message): super.__init__(self, message) async def some_coro(simu_exception=False): delay = round(random.uniform(1.0, 5.0), 2) # We will simulate throwing an exception if simu_exception is True if delay > 4 and simu_exception: raise AsyncException("something wrong!") await asyncio.sleep(delay) return delay ``` Immediately afterward, we will create 2000 concurrent tasks and then use `tqdm_asyncio.gather` instead of the familiar `asyncio.gather` method to see if the progress bar works properly: ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) await tqdm_asyncio.gather(*tasks) print(f"All tasks done.") if __name__ == "__main__": asyncio.run(main()) ``` ![The effect of tqdm_asyncio.gather.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_x-354LEafX4ECRrykXwXMg.gif) The effect of tqdm\_asyncio.gather. Image by Author Ta-da! I finally know where my task is done. Pretty cool. Or let’s replace `tqdm_asyncio.gather` with `tqdm_asyncio.as_completed` and try again: ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) for done in tqdm_asyncio.as_completed(tasks): await done print(f"The tqdm_asyncio.as_completed also works fine.") if __name__ == "__main__": asyncio.run(main()) ``` ![tqdm_asyncio.as_completed also works fine.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_hdYEqdpJ8SJ54AUcIr-sQg.gif) tqdm\_asyncio.as\_completed also works fine. Image by Author Great, it still works fine. --- ## Advanced Tips and Tricks ### Some common configuration items tqdm has a rich set of [configuration items](https://github.com/tqdm/tqdm?ref=dataleadsfuture.com#parameters), so here are some common ones. - `desc`. You can configure a desc parameter to display a title in front of the progress bar, which is useful when distinguishing between multiple groups of tasks. ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) _ = await tqdm_asyncio.gather(*tasks, desc="The progress of works") print(f"The role of the desc configuration item.") if __name__ == "__main__": asyncio.run(main()) ``` ![The role of the desc configuration item.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_5lDhV2I7vv8YAAlrUB_WWQ.gif) The role of the desc configuration item. Image by Author - `ncols`. If the default progress bar is too short, you can make it longer with this parameter. ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) _ = await tqdm_asyncio.gather(*tasks, desc="The progress of works", ncols=105) print(f"Use ncols to change the width of the bar.") if __name__ == "__main__": asyncio.run(main()) ``` ![Use ncols to change the width of the bar.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_od3v9ib0MdisWqG1b-HEUw.gif) Use ncols to change the width of the bar. Image by Author - `colour`. Pycharm’s cli shows the progress bar in red by default, which is still too harsh, so you can use this parameter to change the bar to another color. But as of writing this article, I still haven’t found a way to change the text to white. ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) _ = await tqdm_asyncio.gather(*tasks, colour="white") print("Use colour to change the color of the bar.") if __name__ == "__main__": asyncio.run(main()) ``` ![Use colour to change the color of the bar.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_5NzOMwbV2NnKZj4I4TwNqg.gif) Use colour to change the color of the bar. Image by Author - `bar_format`. This option allows you to flexibly control the content and format of the progress bar display. For example, if you want to display an ETA at the top. ```Python async def main(): tasks = [] for _ in range(2000): tasks.append(some_coro()) _ = await tqdm_asyncio.gather(*tasks, bar_format="Eta:{eta}.|{bar}{r_bar}") print("Use bar_format to customize the content of the progress bar.") if __name__ == "__main__": asyncio.run(main()) ``` ![Use bar_format to customize the content of the progress bar.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_xbQ0CZeptkEv1mShgck7vA.gif) Use bar\_format to customize the content of the progress bar. Image by Author ### Handling of exceptions As you can see from the source code, tqdm implements the `gather` method through the `tqdm_asyncio.as_completed` method. Therefore, we can’t skip exception-catching by using the `return_exceptions` parameter. Which is a pity. But we can still handle exceptions within `tqdm_asyncio.as_completed` via `try…exception` in `tqdm_asyncio.as_completed`: ```Python async def main(): tasks, errs = [], 0 for _ in range(2000): tasks.append(some_coro(simu_exception=True)) for done in tqdm_asyncio.as_completed(tasks): try: _ = await done except AsyncException: errs += 1 print(f"All tasks done. {errs} task(s) failed") if __name__ == "__main__": asyncio.run(main()) ``` ![Handling of exceptions.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/1_pHwS080TFMBdsglBEd9cNg.gif) Handling of exceptions. Image by Author --- ## Real-World Use Cases Many code examples for asyncio are used `asyncio.sleep` to simulate IO-bound cases, which unfortunately oversimplifies the real-world matter. We should use a real-world case to explain using tqdm in asyncio. However, we cannot use a real case in this chapter for space reasons. In the next chapter, we will demonstrate how the tqdm progress bar works in the real world in an example of using asyncio to implement a map-reduce program to handle large files. [Combining Multiprocessing and Asyncio in Python for Performance BoostsCombining Multiprocessing and asyncio via run\_in\_executor unifies the API for concurrent and parallel programming, simplifies our programming process, and allows us to obtain execution results in order of completion. This article will use a Real-world Example to Explain the Code Implementation![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/0_iQIahHdX4TP6KpbO.jpg)](https://www.dataleadsfuture.com/combining-multiprocessing-and-asyncio-in-python-for-performance-boosts/) --- # Conclusion Using tqdm to indicate progress in asyncio code has many benefits: - We can show progress in the caller’s progress bar without intruding into the business code. - All work can be done in the main process without worrying about thread safety and performance issues. - The graphical presentation is always much more vivid than boring text descriptions. - And all this with just one line of code. I’ve also tried other libraries for progress bars, such as [alive-progress](https://pypi.org/project/alive-progress/?ref=dataleadsfuture.com), which is much cooler in presentation, but alive-progress doesn’t support asyncio. tqdm can also produce some cool effects if set up correctly, but I haven’t delved into it due to time, so feel free to discuss further and leave comments. You may help more interested readers. ### Use These Methods to Make Your Python Concurrent Tasks Perform Better URL: https://www.dataleadsfuture.com/use-these-methods-to-make-your-python-concurrent-tasks-perform-better/ Last updated: 2025-03-26T02:06:32.000Z ## Where the Problem Lies It has always been the case that Python’s multi-threaded performance has never lived up to expectations because of [GIL](https://wiki.python.org/moin/GlobalInterpreterLock?ref=dataleadsfuture.com). So since version 3.4, Python has introduced the asyncio package to execute IO-bound tasks through concurrency concurrently. After several iterations, the asyncio APIs have worked very well, and the performance of concurrent tasks has improved dramatically compared to the multi-threaded version. However, there are still many mistakes that programmers make when using asyncio: One mistake, as shown in the figure below, is to use the await coroutine method directly in a way that changes the call to a concurrent task from asynchronous to synchronous, ultimately losing the concurrency feature. ```Python async def main(): result_1 = await some_coro("name-1") result_2 = await some_coro("name-2") ``` Another mistake is shown in the figure below, although the programmer realizes that he needs to use `create_task` to create a task to be executed in the background. However, the following way of waiting for tasks one by one turns the tasks with different timings into an orderly wait. ```Python async def main(): task_1 = asyncio.create_task(some_coro("name-1")) task_2 = asyncio.create_task(some_coro("name-2")) result_1 = await task_1 result_2 = await task_2 ``` This code will wait for *task\_1* to finish first, regardless of whether *task\_2* finishes first. --- ## What is Concurrent Task Execution So, what is a real concurrent task? Let’s use a diagram to illustrate: ![No matter how many tasks we spawn, we will eventually need to join back.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-6.png) No matter how many tasks we spawn, we will eventually need to join back. Image by Author As the diagram shows, a concurrent process should consist of two parts: starting the background task, rejoining the background task back to the main function, and getting the result. Most readers will already know how to use `create_task` to start a background task. Today, I will introduce a few ways to wait for a background task to complete and the best practices for each. --- ## Getting Started Before we start introducing today’s main character, we need to prepare a sample async method to simulate an IO-bound method call, as well as a custom AsyncException that can be used to kindly prompt an exception message when the test throws an exception: ```Python from random import random, randint import asyncio class AsyncException(Exception): def __init__(self, message, *args, **kwargs): self.message = message super(*args, **kwargs) def __str__(self): return self.message async def some_coro(name): print(f"Coroutine {name} begin to run") value = random() delay = randint(1, 4) await asyncio.sleep(delay) if value > 0.5: raise AsyncException(f"Something bad happen after delay {delay} second(s)") print(f"Coro {name} is Done. with delay {delay} second(s)") return value ``` --- ## Comparison of Methods For Concurrent Execution Once we have done the preparations, it’s time to start the day’s journey and fasten your seat belt. ### [1\. asyncio.gather](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#asyncio.gather) `asyncio.gather` can be used to start a set of background tasks, wait for them to finish executing, and get a list of results: ```Python async def main(): aws, results = [], [] for i in range(3): aws.append(asyncio.create_task(some_coro(f'name-{i}'))) results = await asyncio.gather(*aws) # need to unpack the list for result in results: print(f">got : {result}") asyncio.run(main()) ``` `asyncio.gather`, although it forms a group of background tasks, cannot accept a list or collection as an argument directly. If you need to pass in a list containing background tasks, please unpack it. `asyncio.gather` takes a `return_exceptions` argument. When the value of `return_exception` is False, if any background task throws an exception, it will be thrown to the caller of the gather method. And the result list of the gather method is empty. ```Python async def main(): aws, results = [], [] for i in range(3): aws.append(asyncio.create_task(some_coro(f'name-{i}'))) try: results = await asyncio.gather(*aws, return_exceptions=False) # need to unpack the list except AsyncException as e: print(e) for result in results: print(f">got : {result}") asyncio.run(main()) ``` ![Exception catching of asyncio.gather with return_exceptions=False.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-7.png) Exception catching of asyncio.gather with return\_exceptions=False. Image by Author When the value of `return_exception` is True, exceptions thrown by background tasks will not affect the execution of other tasks and will eventually be merged into the result list and returned together. ```Python results = await asyncio.gather(*aws, return_exceptions=True) ``` ![Exception catching of asyncio.gather with return_exceptions=True.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-8.png) Exception catching of asyncio.gather with return\_exceptions=True. Image by Author Next, let’s see why the `gather` method can’t accept a list directly, but has to unpack the list. Because when a list is filled and executed, it is difficult to add new tasks to the list while we wait for them to finish. However, the `gather` method can use nested groups to mix existing tasks with new ones, which solves the problem of not being able to add new tasks in the middle: ```Python async def main(): aws, results = [], [] for i in range(3): aws.append(asyncio.create_task(some_coro(f'name-{i}'))) group_1 = asyncio.gather(*aws) # note we don't use await now # when some situation happen, we may add a new task group_2 = asyncio.gather(group_1, asyncio.create_task(some_coro("a new task"))) results = await group_2 for result in results: print(f">got : {result}") asyncio.run(main()) ``` However, `gather` cannot set the timeout parameter directly. If you need to set a timeout for all running tasks, use this pose, which is not elegant enough. ```Python async def main(): aws, results = [], [] for i in range(3): aws.append(asyncio.create_task(some_coro(f'name-{i}'))) results = await asyncio.wait_for(asyncio.gather(*aws), timeout=2) for result in results: print(f">got : {result}") asyncio.run(main()) ``` ### [2\. asyncio.as\_completed](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#asyncio.as%5Fcompleted) Sometimes, we must start the following action immediately after completing one background task. For example, when we crawl some data and immediately call the machine learning model for computation, the `gather` method cannot meet our needs, but we can use the `as_completed` method. Before using `asyncio.as_completed` method, let’s look at this method’s source code. ```Python # This is *not* a @coroutine! It is just an iterator (yielding Futures). def as_completed(fs, *, timeout=None): # ... for f in todo: f.add_done_callback(_on_completion) if todo and timeout is not None: timeout_handle = loop.call_later(timeout, _on_timeout) for _ in range(len(todo)): yield _wait_for_one() ``` The source code shows that `as_completed` is not a concurrent method, and returns an iterator with a `yield` statement. So we can directly iterate over each completed background task, and we can handle exceptions for each task individually without affecting the execution of other tasks: ```Python async def main(): aws = [] for i in range(5): aws.append(asyncio.create_task(some_coro(f"name-{i}"))) for done in asyncio.as_completed(aws): # we don't need to unpack the list try: result = await done print(f">got : {result}") except AsyncException as e: print(e) asyncio.run(main()) ``` `as_completed` accepts the `timeout` argument, and the currently iterated task after the timeout occurs will throw `asyncio.TimeoutError`: ```Python async def main(): aws = [] for i in range(5): aws.append(asyncio.create_task(some_coro(f"name-{i}"))) for done in asyncio.as_completed(aws, timeout=2): # we don't need to unpack the list try: result = await done print(f">got : {result}") except AsyncException as e: print(e) except asyncio.TimeoutError: # we need to handle the TimeoutError print("time out.") asyncio.run(main()) ``` ![The result of running asyncio.as_completed with timeout parameter.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-9.png) The result of running asyncio.as\_completed with timeout parameter. Image by Author `as_completed`is much more flexible than `gather` in terms of handling the results of task execution, but it is difficult to add new tasks to the original task list while waiting. ### [3\. asyncio.wait](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#asyncio.wait) `asyncio.wait` is called in the same way as `as_completed`, but returns a tuple with two sets: `done` and `pending`. `done` holds the tasks that have finished executed, and `pending` holds the still-running tasks. `asyncio.wait` accepts a `return_when` parameter, which can take three enumerated values: - When `return_when` is `asyncio.ALL_COMPLETED`, `done` stores all completed tasks, and `pending` is empty. - When `return_when` is `asyncio.FIRST_COMPLETED`, `done` holds all completed tasks, and `pending` holds the still-running tasks. ```Python async def main(): aws = set() for i in range(5): aws.add(asyncio.create_task(some_coro(f"name-{i}"))) done, pending = await asyncio.wait(aws, return_when=asyncio.FIRST_COMPLETED) for task in done: try: result = await task print(f">got : {result}") except AsyncException as e: print(e) print(f"the length of pending is {len(pending)}") asyncio.run(main()) ``` ![The result of running asyncio.wait.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-10.png) The result of running asyncio.wait. Image by Author - When `return_when` is `asyncio.FIRST_EXCEPTION`, `done` stores the tasks that have thrown exceptions and completed execution, and `pending` holds the still-running tasks. When `return_when` is `asyncio.FIRST_COMPLETED `or `asyncio.FIRST_EXECEPTION`, we can call `asyncio.wait `recursively, so that we can add new tasks and keep waiting for all tasks to finish, depending on the situation. ```Python async def main(): pending = set() for i in range(5): pending.add(asyncio.create_task(some_coro(f"name-{i}"))) # note the type and name of the task list while pending: done, pending = await asyncio.wait(pending, return_when=asyncio.FIRST_EXCEPTION) for task in done: try: result = await task print(f">got : {result}") except AsyncException as e: print(e) pending.add(asyncio.create_task(some_coro("a new task"))) print(f"the length of pending is {len(pending)}") asyncio.run(main()) ``` ![we can call asyncio.wait recursively When return_when is asyncio.FIRST_COMPLETED or asyncio.FIRST_EXECEPTION.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-11.png) we can call asyncio.wait recursively When return\_when is asyncio.FIRST\_COMPLETED or asyncio.FIRST\_EXECEPTION. Image by Author ### [4\. asyncio.TaskGroup](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#task-groups) In Python 3.11, asyncio introduced the new `TaskGroup` API, which officially enables Python to support [**Structured Concurrency**](https://en.wikipedia.org/wiki/Structured%5Fconcurrency?ref=dataleadsfuture.com). This feature allows you to manage the life cycle of concurrent tasks in a more Pythonic way. For the sake of space, I won’t go into too much detail here, but interested readers can refer to my article: [Why Taskgroup and Timeout Are so Crucial in Python 3.11 AsyncioEmbracing Structured Concurrency in Python 3.11![](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/size/w256h256/2023/08/color-4.png)Data Leads FuturePeng Qian![](https://images.unsplash.com/photo-1603880921125-88ce2fc04673?crop=entropy&cs=tinysrgb&fit=max&fm=jpg&ixid=M3wxMTc3M3wwfDF8c2VhcmNofDV8fGNoYW9zfGVufDB8fHx8MTY5MTg4ODQ1NHww&ixlib=rb-4.0.3&q=80&w=2000)](https://www.dataleadsfuture.com/why-taskgroup-and-timeout-are-so-crucial-in-python-3-11-asyncio/) --- ## Conclusion This article introduced the `asyncio.gather`, `asyncio.as_completed`, and `asyncio.wait` APIs, and also reviewed the new `asyncio.TaskGroup` feature introduced in Python 3.11. Using these background task management methods according to actual needs can make our asyncio concurrent programming more flexible. Due to experience, there are inevitably omissions in the exposition of this article, so please feel free to leave comments during the reading process, and I will reply actively. ### Why Taskgroup and Timeout Are so Crucial in Python 3.11 Asyncio URL: https://www.dataleadsfuture.com/why-taskgroup-and-timeout-are-so-crucial-in-python-3-11-asyncio/ Last updated: 2025-03-26T02:06:09.000Z In last year’s Python 3.11 release, the asyncio package added the[TaskGroup](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#task-groups)` `and [timeout](https://docs.python.org/3/library/asyncio-task.html?ref=dataleadsfuture.com#asyncio.timeout)APIs. These two APIs introduced the official [Structured Concurrency](https://en.wikipedia.org/wiki/Structured%5Fconcurrency?ref=dataleadsfuture.com) feature to help us better manage the life cycle of concurrent tasks. Today, I’ll introduce you to using these two APIs and the significant improvements Python has brought to our concurrent programming with the introduction of Structured Concurrency. --- ## New Features of The Python 3.11 Asyncio Package ### TaskGroup `TaskGroup `is created using an asynchronous context manager, and concurrent tasks can be added to the group by the method `create_task`, with the following code example: ```Python async def main(): async with asyncio.TaskGroup() as tg: tg.create_task(some_coro(1)) tg.create_task(other_coro(2)) print("Both tasks have completed now.") ``` When the context manager exits, it waits for all tasks in the group to complete. While waiting, we can still add new tasks to `TaskGroup`. Note that assuming that a task in the group throws an exception other than `asyncio.CancelledError` while waiting, all other tasks in the group will be canceled. Also, all exceptions were thrown except for `asyncio.CanceledError` will be combined and thrown in the [ExceptionGroup](https://docs.python.org/3/library/exceptions.html?ref=dataleadsfuture.com#exception-groups). ### Timeout `asyncio.timeout` is also created using the asynchronous context manager. It limits the execution time of concurrent code in a context. Let’s assume that if we need to set a timeout to a single function call, it is sufficient to call `asyncio.wait_for`: ```Python async def main(): await asyncio.wait_for(some_coro(delay=2), timeout=1) ``` But when it is necessary to set a uniform timeout for multiple concurrent calls, things will become problematic. Let’s assume we have two concurrent tasks and want them to run to completion in 8 seconds. Let’s try to assign an average timeout of 4 seconds to each task, with code like the following: ```Python async def main(): await asyncio.wait_for(some_coro(delay=5), timeout=4) await asyncio.wait_for(other_coro(delay=2), timeout=4) ``` You can see that although we set an average timeout for each concurrent method, such a setting may cause uncontrollable situations since each call to the IO-bound task is not guaranteed to return simultaneously, and we still got a `TimeoutError`. At this point, we use the `asyncio.timeout` block to ensure that we set an overall timeout for all concurrent tasks: ```Python async def main(): async with asyncio.timeout(delay=6): async with asyncio.TaskGroup() as tg: tg.create_task(some_coro(delay=5)) tg.create_task(other_coro(delay=2)) print("All tasks have completed in time.") ``` ### What is Structured Concurrency `TaskGroup` and `asyncio.timeout` above uses the `async with` feature. Just like `with` struct block can manage the life cycle of resources uniformly like this: ```Python def main(): with open("hello.txt", "w") as f: f.write("hello world.") ``` But calling concurrent tasks inside `with` block does not work because the concurrent task will continue executing in the background while the `with` block has already exited, which will lead to improper closure of the resource: ```Python async def file_coro(f): await asyncio.sleep(5) f.write("hello world.") async def main(): with open("hello.txt", "w") as f: # This will result in nothing being written to the file asyncio.create_task(file_coro(f)) # Do some other things. ``` Therefore, we introduced the `async with` feature here. As *with*, *async with* and `TaskGroup` Is used to manage the life cycle of concurrent code uniformly, thus making the code clear and saving development time. We call this feature our main character today: [**Structured Concurrency**](https://en.wikipedia.org/wiki/Structured%5Fconcurrency?ref=dataleadsfuture.com). --- ## Why Structured Concurrency Is So Important ### History of concurrent programming Before the advent of concurrent programming, we executed our code serially. Code would perform `for_loop` loops, `if_else` conditional branches, and function calls sequentially, depending on the order in the call stack. ![The run order in different code structures.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-2.png) The run order in different code structures. Image by Author However, as the speed of code execution became more and more demanding in terms of computational efficiency and as computer hardware developed significantly, parallel programming (CPU bound) and concurrent programming (IO bound) gradually emerged. Before coroutine emerged, Python programmers used threading to implement concurrent programming. But Python’s threads have a problem, that is, [GIL (Global Interpreter Lock)](https://towardsdatascience.com/python-gil-e63f18a08c65?ref=dataleadsfuture.com), the existence of GIL makes the thread-based Concurrency unable to achieve the desired performance. So asyncio coroutine emerged. Without GIL and inter-thread switching, concurrent execution is much more efficient. If threads are time-slice-based task switching controlled by the CPU, then coroutine is the creation and switching of subtasks back into the hands of the programmer himself. While programmers enjoy convenience, they also encounter a new set of problems. ### Problems with the concurrent programming model As detailed in [this article](https://vorpus.org/blog/notes-on-structured-concurrency-or-go-statement-considered-harmful/?ref=dataleadsfuture.com), concurrent programming raises several issues regarding control flow. Concurrent programming is opening up multiple branch processes in our main thread. These branch tasks silently perform network requests, file accesses, database queries, and other duties in the background. Concurrent programming will change the flow of our code from this to this: ![Concurrent programming will change the flow of our code](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-3.png) Concurrent programming will change the flow of our code. Image by Author According to the “low coupling, high cohesion” rule of programming, we all want to join all the background tasks in a module together after execution like this: ![We all want to join all the background tasks in a module together after execution.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-4.png) We all want to join all the background tasks in a module together after execution. Image by Author But the fact is that since multiple members develop our application or call numerous third-party components, we need to know which tasks are still executing in the background and which tasks are finished. It’s more likely that one background task will branch into several other branch tasks. Ultimately, these branching tasks need to be found by the caller and wait for their execution to complete, so it becomes like this: ![Although this is not Marvel’s multiverse, the situation is now just like the multiverse.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-5.png) Although this is not [Marvel’s multiverse](https://en.wikipedia.org/wiki/Multiverse%5F%28Marvel%5FComics%29?ref=dataleadsfuture.com), the situation is now just like the multiverse. Image by Author Although this is not [Marvel’s multiverse](https://en.wikipedia.org/wiki/Multiverse%5F%28Marvel%5FComics%29?ref=dataleadsfuture.com), the situation is now just like the multiverse, bringing absolute chaos to our natural world. Some readers say that `asyncio.gather` could be responsible for joining all the background tasks. But `asyncio.gather` it has its problems: - It cannot centrally manage backend tasks in a unified way. Often creating backend tasks in one place and calling `asyncio.gather` in another. - The argument `aws` received by `asyncio.gather` is a fixed list, which means that we have set the number of background tasks when `asyncio.gather` is called, and they cannot be added randomly on the way to waiting. - When a task is waiting in `asyncio.gather` throws an exception, it cannot cancel other tasks that are executing, which may cause some tasks to run indefinitely in the background and the program to die falsely. Therefore, the Structured Concurrency feature introduced in Python 3.11 is an excellent solution to our concurrency problems. It allows the related asynchronous code to all finish executing in the same place, and at the same time, it will enable `tg` instances to be passed as arguments to background tasks, so that new background tasks created in the background tasks will not jump out of the current life cycle management of the asynchronous context. Thus, Structured Concurrency is a revolutionary improvement to Python asyncio. --- ## **Comparison with Other Libraries That Implement Structured Concurrency** Structured Concurrency is not the first of its kind in Python 3.11; we had several concurrency-based packages that implemented this feature nicely before 3.11. **Nurseries in Trio** [Trio](https://trio.readthedocs.io/en/stable/?ref=dataleadsfuture.com) was the first library to propose Structure Concurrency in the Python world, and in Trio, the API `open_nursery` is used to achieve the goal: ```Python async def main(): async with trio.open_nursery() as nursery: nursery.start_soon(child1, 1) nursery.start_soon(child2, 2) print("All tasks done.") ``` ### create\_task\_group in Anyio But with the advent of the official Python asyncio package, more and more third-party packages are using asyncio to implement concurrent programming. At this point, using Trio will inevitably run into compatibility problems. At this point, [Anyio](https://anyio.readthedocs.io/en/stable/index.html?ref=dataleadsfuture.com), which claims to be compatible with both asyncio and Trio, emerged. It can also implement Structured Concurrency through the `create_task_group` API: ```Python import anyio async def some_task(num: int = 0): print(f"Task {num} running") await anyio.sleep(num) print(f"Task {num} finished") async def main(): async with anyio.create_task_group() as tg: for num in range(5): tg.start_soon(some_task, num) print("All tasks finished!") anyio.run(main) ``` ### Using quattro in low Python versions If you want to keep your code native to Python to easily enjoy the convenience of Python 3.11 asyncio in the future, there is a good alternative, [quattro](https://github.com/Tinche/quattro?ref=dataleadsfuture.com), which has fewer stars and is risk-averse. --- ## Conclusion The TaskGroup and timeout APIs introduced in Python 3.11 bring us the official Structured Concurrency feature. With Structured Concurrency, we can make concurrent programming code better abstracted, and programmers can more easily control the life cycle of background tasks, thus improving programming efficiency and avoiding errors. Because of limited experience, if there are any omissions in this article about concurrent programming or Structured Concurrency, or if you have better suggestions, please comment. I will be grateful to answer you. ### Implement a Cache Decorator with Time to Live Feature in Python URL: https://www.dataleadsfuture.com/implement-a-cache-decorator-with-time-to-live-feature-in-python/ Last updated: 2025-04-08T01:04:52.000Z ## The Problem The [lru\_cache](https://docs.python.org/3/library/functools.html?ref=dataleadsfuture.com#functools.lru%5Fcache) decorator in Python’s functools package provides an implementation based on an LRU cache. Using this decorator functions with the same arguments will be significantly faster from the second time they are executed. However, lru\_cache cannot support cache expiration. If you want the cache to expire after a certain amount of time to update the cache when the function is called next time, lru\_cache cannot achieve this. --- ## How to Solve ### 1\. Implement a lru\_cache with a TTL feature Therefore, I have implemented a new decorator based on lru\_cache. This decorator can accept a ttl parameter. This parameter can accept a time in seconds, and when this time expires, the next function call will return a new value and refresh the cache. For those who need to solve the problem urgently, here is the source code: ```Python from functools import lru_cache, update_wrapper from typing import Callable, Any from math import floor import time def ttl_cache(maxsize: int = 128, typed: bool = False, ttl: int = -1): if ttl <= 0: ttl = 65536 hash_gen = _ttl_hash_gen(ttl) def wrapper(func: Callable) -> Callable: @lru_cache(maxsize, typed) def ttl_func(ttl_hash, *args, **kwargs): return func(*args, **kwargs) def wrapped(*args, **kwargs) -> Any: th = next(hash_gen) return ttl_func(th, *args, **kwargs) return update_wrapper(wrapped, func) return wrapper def _ttl_hash_gen(seconds: int): start_time = time.time() while True: yield floor((time.time() - start_time) / seconds) ``` The usage is very simple, like this: ```Python @ttl_cache(maxsize=128, ttl=40) def total_count(n): result = 0 for _ in range(n): result += n return result ``` ### 2\. Test the effectiveness ```Python if __name__ == "__main__": start = time.perf_counter() for i in range(10): time.sleep(6) start = time.perf_counter() print(f"round {i} got {total_count(65736222)} for {time.perf_counter() - start:.2f} seconds.") end = time.perf_counter() print(f"program runs at {end - start:.2f} seconds") ``` We use a ttl\_cache with an expiration time of 40 seconds. Then let the function be executed for 10 rounds, each round for 6 seconds. ![The result after we add ttl_cache decorator to our functions.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image.png) The result after we add ttl\_cache decorator to our functions. Image by Author. As we can see, when it reaches the 7th round, 6\*7=42 (seconds). The cache is refreshed, indicating that our ttl\_cache is successful, hurray. --- ## Some Background Knowledge ### 1\. What is LRU cache? Assuming that the size of the cache is fixed, if the cache is full, some content needs to be deleted to make room for new content. But the question is, which content should be deleted? We certainly hope to delete those caches that are not useful and continue to leave useful data in the cache for later use. So, what data do we consider “useful” data? The LRU cache strategy believes that the data that has been used recently should be “useful”, and the data that has not been used for a long time should be useless. When the capacity is full, those useless data should be deleted first. ### 2\. How is the LRU cache implemented? When implementing LRU, we need to focus on its read-and-write performance. At this point, it is easy to think of using HashMap, which can achieve O(1) speed by accessing data according to the key. But the speed of updating the cache cannot reach O(1) because it needs to determine which data has been accessed the earliest, which requires traversing all the cache to find it. Therefore, we need a data structure that both sorts by access time and can be randomly accessed in constant time. This can be achieved by using HashMap+Doubly linked list. HashMap guarantees O(1) access time for data accessed through the key, and the doubly linked list passes through each data in the order of access time. The reason for choosing a doubly linked list instead of a singly linked list is that it can modify the linked list structure from any node in the middle, without having to traverse from the head node. As shown in the figure below, the black part is the structure of HashMap, and the red arrow is the forward connection of the doubly linked list. It can be seen clearly that the data access order is 1->3->5->6->10\. We only need to change the connection order of the linked list after each access to achieve our goal. ![The lookup order of the LRU cache.](https://storage.ghost.io/c/33/67/33678c00-2c15-4961-93e9-497b427e2006/content/images/2023/08/image-1.png) The lookup order of the LRU cache. Image by Author ### 3\. How does this help us implement ttl\_cache? We know that the LRU algorithm uses HashMap to implement fast data reading, so we can change the parameter by changing the hash key after the expiration time to implement cache expiration. Since the hash key of lru\_cache is calculated based on all hashable parameters in the decorated function, we only need to add a ttl\_hash parameter and change the value of this parameter after the expiration time. --- ## Code Interpretation ### 1\. A general decorator template When starting to write a decorator, I will use a decorator template code, which can help me write a decorator more quickly. For example: ```Python def ttl_cache(maxsize: int = 128, typed: bool = False, ttl: int = -1): def wrapper(func: Callable) -> Callable: def wrapped(*args, **kwargs) -> Any: return func(*args, **kwargs) return update_wrapper(wrapped, func) return wrapper ``` ### 2\. A lazy execution generator code Next, we need to generate a hash key every time the function decorated by ttl\_cache is called, and if it has not exceeded the expiration time, this hash key should remain unchanged. If it exceeds the expiration time, this hash key should be different from the previously generated one. Therefore, I plan to subtract the time when the code starts running from the current time, and then divide the result by the ttl parameter. Finally, the remainder is used as this hash key. Since I need to save time when the code starts running and only get the latest hash key when necessary, the best way I can think of is to use a [generator function](https://wiki.python.org/moin/Generators?ref=dataleadsfuture.com): ```Python def _ttl_hash_gen(seconds: int): start_time = time.time() while True: yield floor((time.time() - start_time) / seconds) ``` In this way, I can use the next() function to get the latest hash key when needed. ### 3\. Use lru\_cache to decorate the original function Next, we will use lru\_cache to decorate the original function, and the new function needs to pass in a ttl\_hash parameter to generate a new hash after the expiration time: ```Python @lru_cache(maxsize, typed) def ttl_func(ttl_hash, *args, **kwargs): return func(*args, **kwargs) ``` In the wrapped function, we get a new hash\_key every time and return a function call with ttl\_hash to achieve our goal. ### 4\. Don’t forget to copy the properties of the original function Finally, don’t forget to copy the module, name, doc, and other attributes of the original function into the newly generated function. The easiest way here is to use update\_wrapper in the functools module: ```Python return update_wrapper(wrapped, func) ``` --- ## Conclusion There are other solutions, such as [expiringdict](https://pypi.org/project/expiringdict/?ref=dataleadsfuture.com). But these solutions change the way how we use lru\_cache, after all, we just want to make lru\_cache more powerful to use. The ttl\_cache implemented in this article is thread-safe and can be used in a multi-threaded environment. I hope you can like my implementation of the TTL cache. And welcome everyone to comment and provide valuable suggestions for improvement. Thank you for reading.