Causal representation learning aims to infer a small set of causally related latent variables and to model their causal relationships in order to explain high-dimensional observed data. Recent studies have made progress in identifying causal representations from time series by hypothesizing temporal structures among latent components. However, these methods are constrained by the observation scale and sampling regularity, as they typically assume that latent components evolve at discrete time steps. Moreover, they do not explicitly identify the causal graph among latent components, which limits the interpretability of the learned representations. To address these limitations, we propose Continuous Causal Component and Structure Discovery (C3SD). We theoretically show that latent causal components and their causal relationships can be identified up to permutation equivalence by modeling synchronous sparsity in the mapping between latent components and observed variables. Building on this result, C3SD employs a dual sparsity-induced autoencoder to infer latent causal components, together with an adaptive group lasso to jointly structure the encoding and decoding matrices. In addition, a neural ordinary differential equation–based joint autoencoder models the continuous-time causal dynamics of the latent components and recovers their underlying dynamical causal structure. Extensive experiments demonstrate that C3SD effectively identifies latent components and the causal mechanisms driving their continuous temporal evolution, particularly in sparsely and irregularly sampled time series.
Dezhi Yang, Jun Wang, C. Domeniconi et al.· Proceedings of the 32nd ACM...· 0 citations
Torus networks are deployed in production AI training clusters for their path diversity and low latency, but 2D Torus scales poorly: electrical packet switches compromise latency, and high-dimensional Torus introduces excessive routing complexity. We present STON (Scalable TOrus Network), a hierarchical architecture that treats a 2D Torus as a supernode and interconnects supernodes with a reconfigurable Optical Circuit Switch (OCS) for AlltoAll-dominated large-scale training networks. STON comprises three coordinated modules: (1) fragmentaware task placement, which minimizes inter-supernode traffic by reducing job fragmentation; (2) non-disruptive logical topology mapping, governed by two principles that prevent OCS reconfiguration from disrupting running tasks or partitioning multisupernode jobs; and (3) compute-phase traffic forwarding, which ensures reachability when direct OCS circuits are unavailable. STON reduces average FCT by 42.2%-61.1% across synthetic workloads and by 52.6% on a one-day Kalos production trace (under an AlltoAll traffic model for all jobs), with 95th-percentile tail latency reduced by up to 74.5%, versus a static direct-connect baseline using the same OCS hardware.
Qinwei Yang, Peirui Cao, Ruyi Zhang et al.· Fall Joint Computer Conferen...· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.