This dataset contains anonymized responses from a questionnaire investigating higher education students' attitudes toward the use of Generative Artificial Intelligence (GenAI) in learning. The survey examines students' perceptions, experiences, and acceptance of GenAI technologies as learning support tools, together with their views on personalization, engagement, learning effectiveness, decision-making support, ethical concerns, and future educational applications. The questionnaire was administered to university students enrolled in higher education programs. All responses were collected anonymously and contain no personally identifiable information. The dataset includes: - demographic characteristics of respondents (e.g., gender, age group, field of study, academic level); - previous experience with Generative AI tools; - frequency and purpose of GenAI use in learning; - Likert-scale responses measuring attitudes toward GenAI-assisted learning; - perceived educational benefits and challenges; - perceptions of personalization, engagement, and decision-support capabilities; - concerns related to academic integrity, privacy, bias, and over-reliance on AI; - coding rules for questionnaire variables. The dataset is intended to support research on artificial intelligence in education, technology acceptance, digital learning, educational decision support, and cognitive computing. It may also be used for statistical analysis, machine learning, educational data mining, structural equation modeling, and comparative studies across institutions or countries.
Semantic search over domain-specific corpora requires an effective embedding model and infrastructure. Elasticsearch’s native semantic_text field and ELSER sparse-vector inference require a commercial Enterprise subscription, inaccessible to most academic institutions. This paper documents the licensing barrier and presents a reproducible manual pipeline achieving comparable semantic search with free Elasticsearch components and open-source small language models (SLMs), on a four-node Raspberry Pi 4 cluster (8 GB RAM, three-node Elasticsearch) over 1088 USGS documents. Five models were evaluated —ELSER v2, .multilingual-e5-small, all-MiniLM-L12-v2, all-mpnet-base-v2, and msmarco-MiniLM-L12-cos-v5—from 35 screened candidates, plus a BM25 lexical baseline. Available process memory, not compute, is the binding constraint: Elasticsearch’s footprint consumes 5–6 GB of the 8 GB. The three 384-dimensional models reindexed the corpus in 2.1–2.2 h; the 768-dimensional all-mpnet-base-v2 took 10.1 h. On-disk size is unreliable for provisioning: .multilingual-e5-small expands from 1.4 GB on disk to 3.5 GB at runtime (2.5×). Retrieval quality was assessed with Precision@10, MRR, MAP@10, and nDCG@10 over 15 human-judged queries rather than raw similarity scores; embedding-based retrieval outperforms BM25, reaching significance for two of five configurations. Msmarco-MiniLM-L12-cos-v5 offers the strongest quality-per-resource trade-off among the 384-dimensional models for 8 GB ARM deployments. Pipeline and configuration artefacts are documented and reproducible.
Stoyan Cheresharov, Georgi Gustinov, Венета Табакова-Комсалова et al.· Electronics· 0 citations
We use cookies to run the site and, with your consent, for analytics and to show ads.
See our Cookie Policy.