Large-scale mining of public transcriptomic datasets can reveal viral diversity that remains invisible to conventional virus-surveillance approaches, while increasingly powerful structure-prediction methods provide a complementary route to characterizing highly divergent viral proteins. Here, we combine sequence detection, phylogenetic analysis, structural prediction, and host-association analyses to investigate the cryptic diversity and biology of a new clade of plant-associated lispi-like viruses. Analyses of RNA-sequencing datasets identified and enabled assembly of 87 coding-complete lispi-like virus genome sequences associated with 75 plant hosts, expanding the known diversity of this new group by approximately 40-fold. The viruses share a conserved four-cistron genome organization, 3′-N(P1)-P2-P3-L(P4)-5′, in which the first cistron is referred to interchangeably as N or P1 and the fourth as L or P4. Structural analyses generate testable structural hypotheses for the four conserved proteins. N (P1) adopts a canonical negative-strand RNA virus nucleocapsid architecture with conserved RNA-interacting residues and a predicted RNA-packaging configuration. P2 is exceptionally divergent, although a subset of structures resembles the ITPase/HAM1 fold. P3 forms a conserved trimeric coiled-coil architecture reminiscent of a viral fusion-protein stalk, but lacks the family-wide sequence features expected of a canonical membrane glycoprotein. L (P4) contains a structurally resolved Mononegavirales-type RNA-dependent RNA polymerase (RdRp) core with invariant catalytic motifs, including the characteristic GDN signature, whereas its accessory regions are substantially more divergent. Phylogenetic insights form a distinct monophyletic lineage sister to the predominantly invertebrate-associated Lispiviridae, supporting their recognition as the new proposed family Masuviridae, comprising 15 tentative genera. Genus-level clustering is accompanied by marked differences in host association, ranging from strong specialization to broader host ranges. Retrospective screening of public sequencing libraries further identified lispi-like virus sequences in 1536 libraries representing 134 plant species, 11 plant families and 182 geographic locations, highlighting a substantial and geographically widespread cryptic virome. Together, these results establish Masuviridae as a deeply divergent lineage of plant-associated negative-sense RNA viruses and illustrate how the integration of sequence, structural, and large-scale transcriptomic analyses can move viral dark matter from detection towards evolutionary and functional characterization.
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
This study investigates how Lean internal startup facilitates software product innovation in large companies and identifies its enablers and inhibitors, and shows the potential of the method-in-action framework to investigate the Lean startup approach in non-startup context.
Henry Edison, Nina M. Smørsgård, Xiaofeng Wang et al.· Journal of Systems and Softw...· 78 citations· ⚡6
This paper highlights the challenges to conduct proper affect-related studies with psychology, provides a comprehensive literature review in affect theory, and proposes guidelines for conducting psychoempirical software engineering.
D. Graziotin, Xiaofeng Wang, P. Abrahamsson· SSE@SIGSOFT FSE· 56 citations· ⚡4
This study conducts a multiple case study on twenty European software startups and proposes a prototype-centric learning model in early stage software startups, and identifies factors that occur as barriers but also facilitators for prototyping in earlystage software startups.
Anh Nguyen-Duc, Xiaofeng Wang, P. Abrahamsson· International Conference on...· 44 citations· ⚡5
It is demonstrated that linker-free PROTACs can outperform traditional designs, marking a paradigm shift in PROTAC development for targeted protein degradation.
Pinal, a 16-billion-parameter foundation model that produces protein candidates from natural-language functional descriptions, supports natural language as a high-level interface for candidate generation in protein design, enabling programmable exploration with reduced reliance on manually specified structural or seque...
A new machine-learning framework aims to improve the success rate of computational protein design while moving away from results that reproduce sequences found in nature.