An experiment using reconstructed scientific-paper writing trajectories reported higher scores across the study’s writing measures than its controls, except on LongBench-Write. The biggest reported differences were 3.5 to 4.0 points on PAW-Bench and 2.4 to 3.9 points on WritingBench’s Academic & Engineering category. On a separate reasoning average, the trajectory conditions scored 45.66 to 46.82, compared with 46.32 after direct supervised fine-tuning and 44.30 after a FineWeb-Edu-only control.
Turning papers into writing records
The study examines whether structured, paper-level writing trajectories are associated with writing capability beyond plain paper text while preserving reasoning and extending long-context reading. Its reverse-construction process makes three products from finished scientific material: whole papers for continued pre-training, or further training on text; passages for supervised fine-tuning; and held-out papers with grading criteria for PAW-Bench.
For the trajectory data, the pipeline reconstructs a writing request, a global plan and a pre-writing deliberation for each section. The paper’s abstract and section text remain verbatim from the source material.
At the scale used for continued pre-training, 1.8 million arXiv papers containing 30 billion source-paper tokens produced 57 billion to 60 billion trajectory tokens per generator. Reported median lengths were 11,000 tokens for source papers and 28,000 to 29,000 tokens for trajectories.
Every run used seed 42. The trajectory condition combined 30 billion unfolded trajectory tokens with 20 billion FineWeb-Edu tokens, for a 50-billion-token total; controls used the same budget with cleaned plain papers or pure FineWeb-Edu. Continued pre-training and supervised fine-tuning used papers dated from 2006 through January 2026, while PAW-Bench drew on a later held-out window from February through June 2026.
The supervised fine-tuning dataset contained 214,500 samples from 148,600 papers across 29 task types. PAW-Bench contained 2,940 tasks, with four tasks for each of 735 selected papers across 15 task types.
The strongest writing result
The highest reported overall writing result came from the Our-Data-4B continued-pre-training condition paired with DeepWriting supervised fine-tuning. It recorded an Average of 54.34, a reported 2.44-point, or 4.7 percent, gain over baseline. The Average gave equal weight to the writing suites after the HelloBench subsets were averaged, and GPT-5.5 served as the common final judge.
A separate comparison reported PAW-Bench moving from 55.42 to 63.73 when 10,000 DeepWriting examples were replaced with writing supervised-fine-tuning samples; the checklist pass rate moved from 0.529 to 0.650. The trajectory-CPT-plus-mixed-SFT condition was reported at 64.13 on PAW-Bench and 54.38 on the overall Average.
The 27B trajectory condition also scored higher than direct supervised fine-tuning on paper-question and long-context tests. Qasper scores were 43.01 versus 39.82, and QASA scores were 24.69 versus 23.46. On LongBench v2, the scores were 43.97 versus 37.93 at up to 32,000 tokens and 40.22 versus 36.96 at up to 64,000 tokens.
Within the trajectory-data family, final training loss was 1.453 for data generated by the 4B model, 1.416 for the 9B model and 1.374 for the 27B model.
What remains uncertain
The reported results are point-score comparisons: the study does not report confidence intervals or repeated-run variability. The writing suites were judged by one language model, GPT-5.5, and the overall Average used equal weighting after the HelloBench subsets were averaged.
The experiments covered one base model and one continued-pre-training mixture: 30 billion trajectory tokens plus 20 billion general-text tokens. The study did not test whether the pattern transfers to larger base models or to other token ratios.
The work is an arXiv preprint, version v1, dated 26 August 2026.
Paper data and sources
Original title: Unfolding Scientific Papers into Multi-Turn Generation Trajectories for Continued Pre-Training
Authors: Qiankai Xu, Qiguang Chen, Zixin Su et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text