The Digital Source For China's Tech Innovation Since 2000

Kuaishou Open-Sources GoLongRL, Overcoming Artificial Intelligence Bottlenecks in Long-Context Reinforcement Learning

June 25, 2026
Editorial Staff

Kuaishou Technology’s large language model team, in collaboration with the University of Chinese Academy of Sciences, has open-sourced GoLongRL, a comprehensive post-training framework designed to solve critical performance degradation in artificial intelligence models processing exceptionally long text sequences.

Current long-context reinforcement learning techniques suffer from highly homogenous training data that focuses almost exclusively on locating specific data points within long essays. This narrow approach leaves models unequipped to handle complex structural text duties like sorting, abstract summaries, or multi-hop logical reasoning. To address this limitation, the Chinese research team released a fully open-source system that includes a high-utility dataset of nearly 23,000 heterogeneous samples, complete training source codes, and a specialized machine learning optimization algorithm.

The dataset features 22,965 precisely cataloged samples structured across nine distinct core task types to comprehensively train long-context comprehension. Rather than relying on synthetic templated inputs, which often teach models to rely on superficial paragraph boundaries, the dataset prioritizes genuine source materials, including literature from Project Gutenberg, academic preprints, legal documents, and corporate financial filings. For domains lacking labeled data, the pipeline synthesizes only the question-and-answer pairs based on raw inputs, ensuring high-fidelity data integrity.

To maximize learning efficiency across varied tasks like sorting, data extraction, and abstract generation, the researchers abandoned traditional single-metric reward functions in favor of localized evaluation scripts. Since text summaries rely on semantic overlaps while ordering sequences depends on ranking coefficients, the system assigns unique evaluation criteria tailored to each task's structural objective.

Managing these varied reward functions introduced complex numerical scaling variances during standard training runs. To stabilize the optimization process, the researchers developed an algorithm named TMN-Reweight. This math-based framework decouples numerical reward scaling from task difficulty corrections, preventing high-variance signals from disrupting model training.

The framework demonstrated immediate performance gains during empirical evaluation. When applied to a small four-billion-parameter base, the data and algorithm configuration surpassed specialized competing long-context models by a notable margin.

Scaling the framework up to a larger thirty-billion-parameter architecture yielded even more substantial results. The resulting model achieved a top evaluation metric score, outperforming elite, closed-source foundation models including DeepSeek-R1, Alibaba’s large-scale Qwen reasoning systems, and Google’s Gemini Flash framework.

Importantly, the reinforcement learning process did not trigger negative capabilities transfer regarding general analytical reasoning. The models demonstrated minor, steady improvements on mainstream intelligence benchmarks while displaying strong capabilities transfer into entirely novel fields, such as agentic memory and multi-turn conversational dialogue recollection.

The framework also displayed significant sequence length extrapolation capabilities. Although the model was trained on a hard maximum limit of 160,000 text tokens, its core synthesis and data retrieval capabilities successfully generalized out to processing blocks containing up to one million text tokens, proving that the learned processing techniques are length-agnostic.

Related Topics: artificial intelligence | Chinese | Chinese Academy of Sciences | data | DeepSeek | education | finance | financial | flash | Kuaishou | language | large language model | legal | literature | LLM | machine learning | performance | research | standard | Tencent | train | training | university

Other News:

Four former ByteDance employees say the company promoted pro-China content to Americans in its now-defunct news app… News

July 27, 2022
upstract.com upstract.com

China, Pacific islands unable to reach consensus on regional pact – Blockchain Tribune

May 31, 2022
blockchaintribune.com blockchaintribune.com

China is giving away digital central bank currency to the people

April 30, 2022
btc-echo.de btc-echo.de

CAE Sells Flight Simulator To China Eastern Airlines

October 28, 2004

EV Sales At Warren Buffett-Backed BYD Tripled In December, Adding To Big Gains By China Makers

January 4, 2022
latestnigeriannews.com latestnigeriannews.com

Knowledge Nugget: Taiwan

February 18, 2025
indianexpress.com indianexpress.com

Trump's Iran Deal Gives Him Nothing He Wanted

April 9, 2026
theatlantic.com theatlantic.com

Global And Chinese Industry Leaders Voice Different Concerns During China's International Anti-Spam Summit

February 1, 2005

Biden and Varadkar discussed ‘challenges posed by China’ and other issues, say US officials

April 17, 2023
irishtimes.com irishtimes.com
  • Contact Us
  • About Us
  • Corrections and Disclosure
  • Privacy Policy
  • Terms & Conditions
  • Contact Us
  • About Us
  • Corrections and Disclosure
  • Privacy Policy
  • Terms & Conditions
© 2026 ChinaTechNews.com. A Service of Asia Media Network.