A new preprint proposes a practical way to describe construction work in terms a robot could use: 41 basic actions, arranged into 12 groups and three broad classes. The system, called TARCAT, is intended to give researchers and engineers a shared vocabulary for the capabilities needed on a construction site.
The study does not show that robots can independently carry out construction jobs. Instead, it builds an occupation-based catalogue from task descriptions and instructional videos, then reports a small hardware demonstration of selected actions. The authors present the result as a starting point for robot skill libraries and capability planning.
From job descriptions to observable movements
The researchers began with O*NET task statements and employment profiles from the Bureau of Labor Statistics, using the profiles to map employment information to O*NET occupations. Their selected set covered seven occupations, including construction laborers, carpenters, electricians, plumbers, painters, brickmasons and roofers.
Across those occupations, the source material contained 169 O*NET task statements. Of these, 137 involved physical movement and 32 did not. The distinction mattered because the project focused on actions that could be observed, described and eventually linked to robot behavior.
For the video portion, the authors used ChatGPT with GPT-5.6 to search YouTube for demonstrations of movement-based work. They kept videos in which the relevant activity was clearly visible and related to the task, and mapped each retained video to every task it demonstrated. The primary corpus contained 30 instructional videos covering 91 O*NET tasks.
Non-movement activities were inferred from the O*NET descriptions. Human annotators marked movement activities with timestamps in the videos, and the activities and their subgoal sequences were assigned primitive and skill labels. That combination gave the taxonomy both a job-level source and examples of what the work looked like in practice.
A building block approach to robot skills
TARCAT treats a skill as an ordered sequence of activities carrying primitive labels, rather than as a description tied to only one particular task. Variables in the sequence can be filled with tools or objects observed in the demonstration. In principle, that lets a common action pattern be represented across different pieces of work without treating every task as a completely new skill.
The task representation also allows actions to happen at the same time, to repeat, or to contain another skill. Simultaneous actions are recorded in order, repeated actions are marked as repetitions, and nested skills are invoked as subskills. These rules are meant to preserve the structure of real work while keeping the underlying action vocabulary reusable.
That structure is the paper’s central engineering proposal. By connecting occupational task statements with observable demonstrations, the authors interpret TARCAT as a human-readable specification of robot capabilities. The representation is intended to organise demonstrations, specify robot requirements, support coding agents and assemble larger skills from smaller ones, although those uses are presented as intended applications rather than results of a comparative test.
A small hardware test
The researchers then selected four tool-use activities for a physical demonstration. They used a six-axis DOBOT CR3 robotic arm fitted with a 15-actuator, tendon-driven CRAFT hand. The authors report that selected elements of the taxonomy could be mapped to executable physical behaviors on this setup.
That result is narrower than a robot performing a construction task from start to finish. The study reports no formal success or failure rate, timing measure, or test of how well the behaviors generalised beyond the scripted demonstrations. It also includes no comparison group or controlled performance evaluation, so it cannot show that the taxonomy improves task success, learning speed or transfer to new situations.
What remains to be tested
The authors describe the hardware work as preliminary and acknowledge that the coverage is limited to seven occupations, with few demonstrations for each task. The reported study does not measure agreement between annotators or how complete the task coverage is. It also does not evaluate long-horizon composition, in which a robot must reliably combine many skills over an extended job.
Future work identified by the paper includes expanding the corpus, measuring annotation agreement and coverage, implementing more primitives on mobile manipulators and humanoid robots, and testing long-horizon skill composition. Those steps would determine whether the proposed vocabulary can move beyond a useful catalogue toward a reliable foundation for construction robots.
The annotations are available through the TARCAT GitHub repository, and the paper also links to a demonstration video. The supplied document is arXiv version 1, dated Aug. 26, 2026, and no journal publication status is reported.
Paper data and sources
Original title: A Taxonomy of Construction Task Activities for Robot Workers
Authors: Sadman Sakib, Zhangyi None Peng, Yujie Pang et al.
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-26
DOI: Not available
Original paper · Full text