Skip to main content

News · Research

EvolveScaler: Tencent and Renmin University Release Information-Evolution Benchmark Framework

A new framework and dataset combining executable state machines with natural-language rendering has been made public to measure how well large language models track evolving information.

2 min read Source: github.com

What changed: Tencent Hunyuan and Renmin University released an executable state machine-based framework and accompanying dataset for measuring information evolution in language models.

Tencent Hunyuan and the Gaoling School of Artificial Intelligence at Renmin University of China have published a GitHub repository and project site for a paper titled 'EvolveScaler: Synthesizing Information-Evolution Contexts via Executable State Machines and Natural-Language Rendering.' The framework is designed to generate training and evaluation data for large language models by simulating how information changes over time. Authors include Ziliang Zhao, Zenan Xu, Shuting Wang, Zhao Wang, Bowen Cao, Minda Hu, Lincheng Li, Pluto Zhou, and Zhicheng Dou; Zhao and Xu contributed equally, while Zhou and Dou are listed as corresponding authors.

EvolveScaler operates in three stages: annotators specify workflows as typed transitions; a large language model compiles those specifications into executable simulators; and the simulators, through replay, produce both natural-language dialogue and deterministically computed reference answers. The accompanying dataset comprises 117 task prototypes across 12 thematic domains and approximately 35,100 training examples, alongside 585 validated evaluation instances whose difficulty ranges from roughly 7 to 1,200 events per instance across five tiers.

Benchmark results expose a sharp performance gap tied to context length: the strongest evaluated model reaches 59.3% avg@5 on the 'very_long' tier, while six models fall below 10% on the longest contexts. An internal A3B model trained on 6,000 examples from the dataset recorded an average gain of 5.25 points across eight independent benchmarks. The repository is publicly available under the Tencent-Hunyuan GitHub organization.

Key facts

  • The framework operates in three stages: workflow specification via typed transitions, LLM-based compilation into executable simulators, and replay to produce natural-language dialogue with deterministic reference answers.
  • The dataset includes 117 task prototypes across 12 thematic domains, approximately 35,100 training examples, and 585 validated evaluation instances.
  • The strongest model achieves 59.3% avg@5 on the 'very_long' evaluation tier; six models score below 10% on the longest contexts.
  • An internal A3B model trained on 6,000 examples gained an average of 5.25 points across eight independent benchmarks.
  • The repository is publicly available under the Tencent-Hunyuan GitHub organization.

Why it matters

The dataset provides concrete, measurable evidence of how sharply model performance degrades with context length, exposing a current limitation in large language models' ability to track evolving information.

Watch next: The publication venue and the identity of the 'A3B' model remain undisclosed; clarification from the authors is expected.

Source date: 2026-09-09T06:21:05.327Z · Verification confidence: 92/100

Related reading

First conversation

Tell us what you want to do, and we will work out together where to start.

In the first call we talk through your business, where things stand and what matters most. We say plainly which parts make sense for us to take on and which you should run yourself.

Cookies and measurement

Apart from what the site needs to work, measurement or advertising tags only run if you allow them. No measurement tags are active on this site right now. Details