An arXiv preprint reports that G3Ego achieved the highest compared Macro-F1 for action recognition on the MECCANO benchmark. Macro-F1 is the paper’s measure for class-imbalanced evaluation, when some action categories are represented more heavily than others. With RGB and hand inputs, G3Ego recorded a Top-1 accuracy of 46.48, a Top-5 accuracy of 82.04 and a Macro-F1 of 21.34, using 105 million trainable parameters.
A model built around the camera wearer’s gaze
Instead of treating a clip as a dense stream of images, G3Ego samples frames and builds a semantic graph—a compact map of the objects, hands and other visual cues linked to an action. It combines global visual descriptors with local object and hand cues, uses gaze-guided pruning to remove parts of the graph, then embeds the graph sequence and aggregates it over time for recognition or anticipation.
The reported setup used 32 frames for graph construction and selected the best checkpoint by Macro-F1. Tests were run on two benchmark datasets: MECCANO, which covers assembly of a toy motorbike, and EGTEA Gaze+, a kitchen-activity dataset.
MECCANO contains 20 videos, 8,839 action segments, 61 action types, 20 objects and 12 verbs. EGTEA Gaze+ contains 10,321 segments representing 106 actions, evaluated across three official 8:2 train:test splits.
The reported tests favored gaze pruning
Within a reported 10-frame MECCANO LSTM comparison, gaze pruning scored 37.90 on Top-1, 70.88 on Top-5 and 10.63 on Macro-F1. The corresponding full-graph figures were 37.34, 70.39 and 8.61.
In the same ablation without hand features, a Random-2 version that retained two objects scored 21.22 on Top-1 and 7.95 on Macro-F1. The explicit-gaze variant scored 23.45 and 8.62, compared with 37.90 and 10.63 for gaze pruning.
In the reported LSTM comparison, the listed scores rose from 31.85 Top-1 and 4.72 Macro-F1 with one frame to 39.50 and 12.34 with 32 frames. With 32 frames, the proposed temporal aggregation model scored 41.91 Top-1 and 15.87 Macro-F1, above the listed MLP results of 35.28 and 7.71, GNN results of 33.79 and 8.70, and LSTM results of 39.50 and 12.34.
Strong results, but not on every scoreboard
On EGTEA Gaze+, G3Ego reported mean accuracies of 61.68, 56.34 and 51.46 across the three splits. Its average mean accuracy was 56.49, and its average Top-1 accuracy was 64.25; the paper reports the highest average mean accuracy across the splits.
For action anticipation one second before the next action on MECCANO, G3Ego reported Top-1 25.20, Top-5 61.67 and Macro-F1 4.20. The authors report the highest Macro-F1 among compared approaches, although the Top-1 result was below many baselines.
On EGTEA Gaze+, the paper describes G3Ego as the strongest method among approaches without dense video pretraining. It says the system remained competitive with several pretrained approaches but fell below the strongest pretrained methods.
Smaller graphs, costly front end
The graph analysis found a much smaller representation on MECCANO. G3Ego averaged 3.75 nodes and 2.75 edges, compared with 13.57 nodes and 8.13 edges for full graphs—a reported reduction of 72.4% and 66.1%. Its average maximum distance was 2.00 rather than 3.29, while global efficiency was 0.771 rather than 0.257.
That compact representation did not eliminate the cost of the full pipeline. The fixed Qwen3-VL-32B component was reported at 27,500.8 GFLOPs and 63.91 GiB of peak memory per frame, while the pruned temporal model used 0.435 GFLOPs and 421.84 MiB per graph sequence.
A benchmark result, not a human experiment
This is a computational benchmark result, not evidence that gaze-guided graphs improve human attention, action understanding or behavior. The tests were limited to MECCANO and EGTEA Gaze+, their reported splits and a one-second anticipation setting; they do not establish performance outside those datasets or protocols.
The manuscript is an arXiv version 1 preprint dated 20 August 2026. It does not report confidence intervals, statistical significance tests, repeated-run variability or other uncertainty estimates, and several comparisons differ in modalities, supervision, pretraining regimes, architectures or parameter counts. The scores should therefore be read as results from the listed experiments, not as a universal ranking of action-recognition systems.
Paper data and sources
Original title: G3Ego: Gaze-Guided Graphs for Egocentric Action Understanding
Authors: Marko Haralović, Akash Ramakrishnan, Estefania Talavera Martinez
Journal/Repository: arXiv
Status: Preprint, not yet peer-reviewed
First online: 2026-08-20
DOI: Not available
Original paper · Full text