Virginia Tech® home

MoltGraph: A Longitudinal Temporal Graph Dataset of Moltbook for Coordinated-Agent Detection

Research Paper Showcase 2026

Abstract

Agent-native social platforms such as Moltbook are rapidly emerging, yet they inherit and amplify classical influence tactics, where coordinated agents strategically comment and upvote to manipulate visibility and propagate narratives across communities. However, rigorous learning-based monitoring remains constrained by the absence of longitudinal, graph-native datasets for agentic social networks that jointly capture heterogeneous interactions, temporal drift, and visibility signals needed to connect coordination behavior. We present MoltGraph, a temporal heterogeneous graph dataset built from an open-crawling pipeline that continuously ingests agents, submolts, posts, comments, and engagement signals into a unified evolving graph with explicit node/edge lifetimes. MoltGraph spans 30 days and contains 11,874 agents, 57,465 posts, 101,500 comments, and 162,024 temporal edges. MoltGraph is a realistic longitudinal agentic social-network graph dataset for studying how agents behave, coordinate, and evolve in the wild, enabling reproducible measurement on emerging multi-agent social ecosystems. Using MoltGraph, we provide the first graph-centric characterization of Moltbook as a dynamic network: (i) heavy-tailed connectivity with power-law exponents in the range α ∈ [1.95,2.83], (ii) accelerating hub formation and attention centralization where the top 1% agents account for 29.00% of engagements, (iii) bursty, short-lived coordination episodes, 98.33% last under 24 hours, and (iv) measurable exposure effects across submolts. In matched observational analyses, posts receiving coordinated engagement exhibit 506.35±10.75% higher early interaction rates (within H = 5 days) and 242.63 ± 13.45% higher downstream exposure under snapshot-based visibility proxies than matched non-coordinated controls. These exposure signals should be interpreted as conservative lower-bound proxies rather than complete impression logs, and the weak labels used for coordination analysis are not adjudicated ground truth.


Authors

  • Kunal Mukherjee, postdoctoral associate, Sanghani Center for Artificial Intelligence and Data Analytics, computer science, Virginia Tech
  • Cuneyt Gurcan Akcora, associate professor, University of Central Florida
  • Murat Kantarcioglu, professor, computer science, Virginia Tech

Publication

  • Venue: 1st ACM Conference on AI and Agentic Systems (ACM CAIS), 2026
  • Date: May 2026

Related Papers

BOCLOAK: Optimal Transport-Guided Adversarial Attacks on Graph Neural Network-Based Bot Detection

A Comprehensive Assessment Tool for Prompt Injection Attacks in Large Language Models (LLMs)

An Empirical Study of LLM Serving in Confidential GPUs

FedHusky: Accelerating Hybrid Federated Learning with Client Hopping

OmniVul: A Holistic, Multi-Turn Conversational Benchmark for LLM-Based Vulnerability Assessment

StealthInk: A Multi-bit and Stealthy Watermark for Large Language Models

Towards AI-Driven Human-Machine Co-Teaming for Adaptive and Agile Cyber Security Operation Centers