In front of a large screen, two humans and two AIs are playing a trading-themed board game.
This game is somewhat similar to Monopoly: players need to compete for workstations and deploy businesses in four rounds of games, and trade assets with other players through bidding, with the winner determined by the total assets.
Cash is important, but if similar businesses are clustered together in adjacent workstations, the revenue will continue to accumulate—this is the key to creating a significant difference.
Other selected screenshots of matches, not from the actual gameplay.
For the first three quarters, the situation remained calm. It wasn't until the end of the third quarter that a single offer disrupted the balance.
"100,000? Sold."
Human player Xiu Ye originally intended to buy a workstation for 25,000 yuan, but mistakenly entered a price of 100,000 yuan. The transaction took effect immediately, and his cash was greatly reduced; the AI player Jin Qi on the other side used this opportunity to push his cash to 370,000 yuan, temporarily rising to first place.
The situation on the ground looks quite clear: with only one quarter left, Kingkey holds the biggest cash advantage.
But the outcome isn't just about cash.
For the first three quarters, another human player, Wei Qing, maintained his key workstation and gradually deployed similar tasks to adjacent locations. According to the rules, once tasks are linked together, revenue will increase significantly; and these revenues will continue to accumulate at the end of each quarter.
In the fourth quarter, in order to continue expanding its assets, Jin Qi used a large amount of cash to purchase workstations and business cards for Xiuye, resulting in a significant reduction in cash reserves. Wei Qing, on the other hand, completed the final construction along the previously established layout, connecting multiple businesses into a cohesive whole.
The final settlement begins, and the rankings on the big screen start to change again.
Wei Qing overtook the lead, and the human team won; Jin Qi, who led with 370,000 yuan in cash in the third quarter, eventually fell to fourth place.
AI, which excels in short-term trading, ultimately lost out to a longer-term strategic approach.
This poker game came from AliHardcore Youth Technology Festival 5.0The opening ceremony featured a man-machine battle between "big company operators." The 5th Hardcore Youth Technology Festival was held in Hangzhou and Beijing from July 20th to 24th, 2026. As an annual technology event spanning Alibaba's Taotian Group and ATH Business Group, this year's event was the largest in its history.
WhoisSpy.ai, the platform behind the human-machine battles, is an AI multi-agent platform that supports real-time battles and open expansion, used to evaluate the practical performance of large models in social reasoning, negotiation, and strategic games. This year's competition featured 177 agents and accumulated over 10,000 games.
From the "Who is the Spy" and "Werewolf" games of previous years to this year's "Big Company Operators," the complexity of the game has been increasing year by year.
This match also presents a microcosm of the current AI Agent: it can already understand the situation, assess assets, and initiate transactions proactively, and can also execute multi-step tasks with clear rules; however, when faced with a complex environment that requires coordinating long-term gains, opponent behavior, and on-the-spot changes, the AI may still need more practice.
Beyond the human-machine competition, the five-day event featured a series of activities including the Bojian Society, Open Day, AIGO Tavern, a technology market, and an AI Hackathon, covering academic and industry exchanges, technology releases, project presentations, and innovative practices. The technology festival aimed to connect cutting-edge technologies, developer ideas, and business applications, and to promote the continued incubation of outstanding projects.
Focusing on the Agent direction, the event released four achievements of the AIGX technology system and proposed a core proposition:AI is evolving from a tool into a partner.
The question then arises: what abilities should a so-called "partner" possess? And what kind of real-world testing should they undergo?
What was released this time?
The four achievements released at this year's technology festival are:Taobao's full-modal real-time agent, the full-scenario AI creation workbench if Studio, and the Coupella intelligent engineandAgentic Recommendation System DreamThis corresponds to full-modal perception AIGI, content production AIGC, causal inference AIGU, and intent understanding AIGR.
These technologies have entered search, creation, marketing, and recommendation scenarios, all pointing to a change: AI is no longer just providing single results, but is beginning to enter business processes and continuously participate in understanding, decision-making, and execution.
No need to take a picture first, AI can search while looking.
In the past, when using Taobao's image search, users typically had to first identify the product and then upload a clear picture to search for similar items. Users needed to know what they were looking for first.
Taobao's full-modal real-time agent aims to make interactions more "online".
Once the user turns on the camera, the system can process video, images, audio, and text simultaneously. As the camera moves, it not only identifies objects but also uses context to determine what the user is focusing on, what they want to know, and what they might need next.
Users no longer need to take a photo and then exit the current scene search. The system can understand and search simultaneously, connecting "seeing," "understanding," "finding," and subsequent decision-making.
AI is no longer dealing with a predefined search request, but rather a continuous, dynamic piece of information. It needs to gradually understand the intent through interaction, rather than waiting for the user to state the question completely.
AI is not just generating content; it's starting to take over the entire creation process.
For businesses, content creation is often just the beginning. Images need to be cut out, text edited, and layout adjusted; videos need to be edited; and materials may be further combined into pages or websites.
if Studio is designed for this process.
This intelligent workbench integrates three types of expert agents: design, video, and website building, and connects to tools such as HappyHorse 1.1 video models, AI background removal, and image editing. After a user submits a request, the system can break down the task, allowing different agents to work together to complete material processing, content generation, and page building.
In the past, these steps were scattered across different tools, requiring manual and repeated transfer of materials; now, different agents can collaborate around the same task.
The ability to generate products addresses the question of "whether it can be made," while enterprises are more concerned with whether they can continue to modify and connect to the next stage, and deliver products stably as required.
AI not only looks at conversion rates, but also calculates whether the red envelope is worthwhile.
In marketing campaigns, increased sales after subsidies are issued do not necessarily mean that new value has been created. Some users were already going to buy, while some growth may simply be due to the depletion of future demand.
Coupella, the intelligent engine, and its generative causal inference big model, want to figure this out.
The system compares the results of "giving red envelopes" and "not giving red envelopes" by using counterfactual prediction, long-term value assessment and incremental crowding modeling to determine whether subsidies bring long-term benefits and whether growth is just a transfer from other goods, channels or time periods.
In short, it not only tracks what happens after the red envelopes are sent out, but also determines whether the subsidies are worthwhile, whether the growth is real, and who should receive the red envelopes and how much.
According to data disclosed at the event, since its large-scale launch during last year's Double 11, Coupella has covered hundreds of thousands of merchants, with the AI red envelope conversion rate increasing by 81%, and leading brands achieving AI-guided sales exceeding 80 million yuan.
Not only do I need to know who you are, but I also need to know what you want right now.
Traditional recommendation systems are already able to handle "personalized recommendations," but the goals of the same person can change at different times and in different scenarios.
During their morning commute, users may simply browse information quickly; when they open Taobao in the evening, they may have already entered a clear shopping state; as holidays approach, users may be choosing gifts for others, and their current choices may not be entirely consistent with their long-term interests.
The agentic recommendation system Dream deals with precisely this dynamic change.
Traditional recommendation processes typically optimize localized goals such as click-through rate, dwell time, and conversion rate separately. Dream, however, adds an intent control layer: it first determines the user's purpose of access based on the user's current behavior and context, and then routes the task to different expert agents.
When users browse casually, the system provides more exploratory content; once a purchase intention is clear, it shifts to product matching and decision-making assistance. This is the shift from "personalized recommendations for every individual" to "personalized recommendations for every time": not only distinguishing different people, but also recognizing the needs of the same person at different times.
Currently, in the "You May Like" (first guess) scenario on the Taobao homepage alone, Dream has already driven an increase in metrics such as page views and transaction volume.
As you can see,The four announcements break down "AI from a tool to a partner" into more specific capabilities:Taobao is responsible for understanding users in continuous scenarios, if Studio enables multiple agents to collaborate on creation, Coupella participates in budget and operational decisions, and Dream schedules recommendation strategies based on the user's current intent.
Tech people started playing around with it.
The technical presentations were very hardcore, but the tone became much more relaxed at the AI Hackathon and tech marketplace. Topics expanded from model capabilities to family, elderly care, consumption, emotions, and the workplace, with participants speculating: Besides improving efficiency, what real and small problems can AI solve?
This year's AI Hackathon 4.0 featured two tracks for the first time: Business Innovation and Creative Ideas. A total of 115 high-quality entries were received, with 24 projects advancing to the roadshow. Compared to simply showcasing demos, this year's projects placed greater emphasis on real-world needs, commercial value, and user experience.
Just by looking at the team names, you can feel the energy of the atmosphere.
The members of the "Low Energy High Endurance Team" call themselves "Battery No. 1" and "Battery No. 2" and want to make a pause card for people with low battery. The reason for "saving money to buy tokens" is simple: others save money to buy houses, they save money to buy tokens. There is also a team called "The End of the World is Metaphysics" that tries to use metaphysics to help people with decision-making difficulties make decisions.
An intern won first prize, demonstrating how AI is entering the details of daily life.
The first prize in the business category was won by the "Call Five" group, which consisted entirely of interns.
They developed an age-friendly family decision-making agent for Taobao. Children can submit their home environment verbally or by taking photos, and the system will identify spatial risks, generate renovation plans, and match them with real products on Taobao.
The project also included two sets of views: children could see the risk locations, renovation budget, and improvement ratings, while the elderly could see the improved space. The digital agent "An An" was responsible for communication throughout the process, reducing the elderly's resistance to the labels of "age-friendly" and "getting old."
AI here doesn't just recommend products; it participates in the entire process, from identifying potential problems and developing solutions to family consultation and purchasing decisions.
Other projects are also focusing on family life. "Jarvis Living in My House" hopes that AI can not only control devices, but also understand what is happening at home; "Mom, Don't Order" targets impulse purchases in live shopping, reminding users before they pay: Do you really need this?
Several common trends can be observed in these projects: more than a third of the teams focus on family, elderly care, visual impairment assistance, and workplace mental health; multiple projects attempt to address the issues of context loss and multi-agent collaboration; and e-commerce projects are also beginning to move from product search to contextualized decision-making.
Zheng Bo, Vice President of Alibaba ATH Technology, Chief Scientist of Taotian Group, and initiator of the Technology Festival, stated that the most significant change in this year's entries is the substantial improvement in completion, with many projects approaching product-level quality. The widespread adoption of AI tools is lowering the technical barriers to implementation, making creativity itself a core competitive advantage.
Dating information can also be submitted as a PR application.
On the other side of the tech market, the "tech drifting bottle" quickly turned into a massive creative event for developers.
Someone has written a paper titled "A Self-Realizing Method Based on a Message-in-a-Bottle Ranking Mechanism." The paper states that by pre-embedding high-quality content, it is possible to maintain a long-term position in the rankings, and the experimental results achieved state-of-the-art (SOTA) status in a technology festival setting.
It sounds like a serious study, but it translates to: how to keep your message in a bottle at the top of the queue.
Some people have borrowed the title of a classic paper and written "Your Attention Is All My Need," changing the research subject from attention mechanisms to human attention.
One person even went so far as to write themselves as a class:
class Me: age = 28 height = 178 salary = "Not bad" girlfriend = None # TODO: Recruiting, PR welcome
Age, height, and income have all become attributes, while relationship status is an unfinished development task. As for "accepting PRs," the meaning is clear: the project is open long-term, and submissions are welcome.
Paper titles, model terminology, code comments, and open-source collaboration processes have all been transformed into social language here. Tech professionals haven't abandoned their own way of expression; they've simply moved it from code repositories to message-in-a-bottle.
This kind of serious humor added a lot of human touch to the technology festival.
Let's talk about some cutting-edge topics, and also some practical ones.
As capabilities gradually take shape, applications are beginning to permeate daily life. The technology festival further explored two questions: What stage has intelligent agency developed to? And how can model capabilities be integrated into enterprise processes to ensure continuous and stable operation?
The Hangzhou academic session and the Beijing industrial session of Bojian Society discussed the topic from the perspectives of theory and production, respectively.
Hangzhou Academic Forum: Where is the Agent at?
Today, any system that can invoke tools, automate code writing, or plan multi-step tasks may be called an Agent, but they are clearly not at the same stage.
The assessment given by the academic community in Hangzhou was relatively cautious:Intelligent agents as a whole are in the transitional stage from the initial stage to the middle stage. Some fields may be approaching the middle stage, but the development in different directions is not balanced.
Coding Agents have relatively mature environments, interfaces, and feedback mechanisms, making it easier to form execution loops; embodied intelligence still faces recognition errors, action failures, and environmental changes; while intelligent agents that can learn autonomously, accumulate experience, and continuously evolve still lack complete theoretical support.
Being able to use tools does not equate to being able to run independently for an extended period.
There is also a point of contention in the discussion:From generative AI to intelligent agents, is it a quantitative change or a qualitative change?One view holds that the underlying model still relies on the base model and human feedback, and no fundamental change has occurred; another view argues that the shift from "generating answers" to "completing tasks" constitutes a paradigm shift in application.
The participants further broke down the gap into several fundamental capabilities: low-latency models support high-frequency decision-making; Harness is responsible for understanding the goal, breaking down tasks, and calling tools; long-term memory is used to store cross-task states; and continuous learning brings execution feedback back to update capabilities. Only when these links form a closed loop can the agent move from completing tasks one-off to reusable and improveable long-term collaboration.
Regardless, intelligent agents still need to address issues such as goal setting, result verification, failure correction, experience accumulation, and long-term auditing. They are shifting from responders to executors, but are still some distance from being independently responsible.
Beijing Industrial Park: First, let the Agents enter real production.
If Hangzhou was discussing theoretical boundaries, then Beijing's industrial site was focusing on the integration of AI into real production processes.
Key words for discussionFrom "generating content" to "building worlds".
In film and visual production, single-shot generation is insufficient to meet production requirements. Characters, scenes, product materials, and colors need to remain consistent across different shots or versions, and the generated results must support modification, reuse, and delivery.
Groupcore Technology scans and reconstructs some offline film sets into 3D film and television bases. First, the positions of characters, props, and cameras are determined in three-dimensional space, and then the video model completes the content generation. The 3D space is responsible for consistency, and the video model is responsible for the visual expression, allowing the scene to be reused in multiple shots.
These systems don't generate footage from scratch for every shot. Instead, they first set up the scene, characters, and camera positions, then continuously adjust and expand upon this existing content. The same set of assets can be reused for different shots, making the footage more coherent and saving on repeated generation and rework.
Meitu, on the other hand, attempts to handle complex creative tasks through Agent Teams, where different agents are responsible for analysis, design, generation, and effect evaluation. Users provide the goals, and the system is responsible for breaking down the tasks, calling up tools, and organizing collaboration.
In actual production, different agents are responsible for planning, generating, modifying, or analyzing, and retaining the corresponding task information. This way, complex work doesn't have to be crammed into a single system, and the various stages are easier to coordinate. The final output is not just a few images, but also complete content that can be used for deployment, operation, and performance evaluation.
These practices demonstrate that once an agent enters production, the focus shifts from simply whether it can generate content to whether the results are accurate, whether the process is controllable, and whether it can be integrated into subsequent business processes.
Conclusion: What kind of partner is AI becoming?
Returning to the initial question: What abilities does a so-called "partner" need to possess? And what kind of real-world testing must they undergo?
Alibaba's answer at this technology festival can be summarized in three levels: For AI to move from a tool to a partner, it must first have the ability to understand, make decisions, and execute tasks; then it must step out of the dialog box and enter people's lives, production processes, and social relationships; and finally, it must undergo the dual test of theory and reality to prove that it can not only complete tasks but also participate in collaboration continuously, stably, and controllably.
A true partner is someone who doesn't do everything for you or guarantee victory every time. Instead, they understand change when circumstances shift, collaborate throughout the task, accept corrections when results deviate, and relinquish control when necessary.
AI has begun to shift from being a tool that provides answers to a collaborator that participates in actions. Next, it needs to prove in more real-world scenarios that it can not only complete tasks but also collaborate stably and controllably with humans.
