Oct 2026· East African Journal of Information Technology
Natural Language Processing TechniquesSpeech and dialogue systems
Abstract
Conversational Natural Language Understanding (NLU) has passed through five distinguishable technological eras: rule-based symbolic parsing, statistical and probabilistic modelling, neural sequence learning, Transformer-based pretraining, and the current large language model regime. Existing accounts of this history tend to foreground modelling advances and treat infrastructure as a supporting detail. This article argues instead that the two tracks are causally entangled: each scientific advance in conversational NLU became production-viable only once a matching infrastructure capability, spanning hardware, data pipelines, training paradigms, and serving architecture, made it deployable at scale, and each infrastructure capability in turn exposed the next scientific bottleneck. The analysis traces a specific bottleneck sequence, from combinatorial rule brittleness through representational discreteness, vanishing gradients, and sequential processing limits, to labelled-data scarcity, showing how each resolution reshaped the engineering problems that remained. A closing synthesis argues that large language models have not simplified conversational architecture; they have relocated its complexity from authoring rules and training classifiers toward orchestration, retrieval grounding, and failure isolation across distributed inference systems. The account draws on systems-architecture reasoning applied to the published research and technical-report record, rather than on a new empirical benchmark.
The results are packaged in the Greenfield Startup Model (GSM), which explains the priority of startups to release the product as quickly as possible, and the need to shorten time-to-market, by speeding up the development through low-precision engineering activities.
Carmine Giardino, Nicolò Paternoster, M. Unterkalmsteiner et al.· IEEE Transactions on Softwar...· 178 citations· ⚡14
Software startup companies develop innovative, software-intensive products within limited timeframes and with few resources, searching for sustainable and scalable business models.
M. Unterkalmsteiner, P. Abrahamsson, Xiaofeng Wang et al.· e-Informatica Software Engin...· 157 citations· ⚡17
This study conducts a case survey study based on the secondary data of the major pivots happened in 49 software startups, and demonstrates that customer need pivot is the most common among all pivot types.
Sohaib Shahid Bajwa, Xiaofeng Wang, Anh Nguyen-Duc et al.· Empirical Software Engineeri...· 127 citations· ⚡15
The comparison of adopter and non-adopter sample reveals three potential adoption inhibitor, security, data privacy, and portability, which underlines the importance of the technical and security perspectives for research investigating the adoption of technology.
Nattakarn Phaphoom, Xiaofeng Wang, S. Samuel et al.· Journal of Systems and Softw...· 111 citations· ⚡8
The ongoing work building a Raspberry Pi cluster consisting of 300 nodes is presented, with potential use cases being an inexpensive and green test bed for cloud computing research and a robust and mobile data center for operating in adverse environments.
P. Abrahamsson, S. Helmer, Nattakarn Phaphoom et al.· IEEE International Conferenc...· 110 citations· ⚡7
The results indicate that software developers are a slightly happy population, but the need for limiting the unhappiness of developers remains, and 219 factors representing causes of unhappiness while developing software are identified.
D. Graziotin, Fabian Fagerholm, Xiaofeng Wang et al.· International Conference on...· 84 citations· ⚡6
Related blog posts
MIT News · Artificial Intelligence· news.mit.eduOct 8, 2026