Summary
19 baccalaureate nursing students rated AI-generated virtual reality simulations on the 18-item UTAUT instrument, each item scored 1 to 5. Each construct was tested against the scale midpoint of 3 with a one-sample t-test.
Effort Expectancy is well above neutral (3.74, 95% CI 3.45 to 4.03, d = 1.23, p = <.0001). Performance Expectancy (3.43, p = .0341) and Facilitating Conditions (3.44, p = .0151) are modestly above it. Social Influence is not distinguishable from neutral (3.08, 95% CI 2.62 to 3.54, p = .7215).
Internal consistency is high for the instrument as a whole (alpha 0.912, 95% CI 0.833 to 0.949). At 19 respondents every interval here is wide, and that is the main thing to carry away.
1. The sample and the data
19 respondents answered all 18 items, with no missing values. Every response is an integer from 1 to 5. The response file carries no participant identifier, no timestamp and no free text, so nothing about it is disclosive on its own.
Two things follow from a sample of 19. Confidence intervals are wide, so small differences between constructs are not distinguishable. And because 18 items are scored on five points, the response distributions are lumpy, which matters for the normality assumption in section 4.
2. Acceptance by construct
The comparison of interest is against the scale midpoint of 3, where a respondent neither agrees nor disagrees. A construct score is the mean of its items, so it stays on the original 1 to 5 scale. Cohen's d is the distance from the midpoint in standard deviations, and its interval is a percentile bootstrap over 10,000 resamples.
| Construct | Items | k | Mean | 95% CI | d | t (df) | p | α |
|---|---|---|---|---|---|---|---|---|
| PE Performance Expectancy | Q1-Q6 | 6 | 3.43 | 3.04 to 3.82 | 0.53 | 2.29 (18) | .0341 | 0.91 |
| SI Social Influence | Q7-Q10 | 4 | 3.08 | 2.62 to 3.54 | 0.08 | 0.36 (18) | .7215 | 0.93 |
| EE Effort Expectancy | Q11-Q15 | 5 | 3.74 | 3.45 to 4.03 | 1.23 | 5.35 (18) | <.0001 | 0.84 |
| FC Facilitating Conditions | Q16-Q18 | 3 | 3.44 | 3.10 to 3.78 | 0.62 | 2.69 (18) | .0151 | 0.70 |
The pattern is more interesting than a single acceptance score would be. Students found the technology easy to use, and that is the strongest finding by some distance. They were less convinced it would improve their performance, and they did not report social encouragement to use it at all. For a technology being introduced into a curriculum, those are different problems with different remedies.
3. Every item
Item-level results are exploratory. Eighteen tests on eighteen items inflates the familywise error rate, so Holm-adjusted p-values are given alongside the unadjusted ones. 5 of 18 items survive the adjustment.
| Con. | Item | Mean | SD | Mdn | Agree | Disagree | p | p Holm |
|---|---|---|---|---|---|---|---|---|
| PE | Using virtual reality increases my chances of solving problems I come across | 3.63 | 0.96 | 4 | 14 | 4 | .0099 | .119 |
| PE | Using virtual reality makes my life easier | 3.47 | 0.90 | 4 | 11 | 2 | .0349 | .314 |
| PE | Using virtual reality enables me to accomplish tasks more quickly | 3.05 | 0.97 | 3 | 8 | 6 | .8158 | 1.000 |
| PE | Using virtual reality increases my productivity | 3.47 | 0.96 | 4 | 11 | 4 | .0462 | .370 |
| PE | I find virtual reality useful for my daily life | 3.32 | 1.06 | 4 | 11 | 5 | .2092 | 1.000 |
| PE | Using virtual reality allows me to take responsibility for my own learning | 3.63 | 1.07 | 4 | 10 | 3 | .0187 | .206 |
| SI | Most people who are important to me encourage me to use virtual reality for learning purposes | 2.89 | 1.15 | 3 | 6 | 9 | .6945 | 1.000 |
| SI | Most people who are important to me use virtual reality for learning purposes | 2.89 | 0.94 | 3 | 5 | 8 | .6301 | 1.000 |
| SI | Most people who are important to me think I should use virtual reality for learning purposes | 3.21 | 1.18 | 3 | 7 | 7 | .4477 | 1.000 |
| SI | Most people who are important to me find it helpful to use virtual reality for learning purposes | 3.32 | 0.89 | 3 | 9 | 4 | .1374 | .962 |
| EE | Learning how to use Virtual Reality is easy for me | 3.84 | 0.69 | 4 | 15 | 1 | <.0001 | <.001 |
| EE | I find virtual reality easy to use | 3.79 | 0.79 | 4 | 15 | 2 | .0004 | .006 |
| EE | I can use virtual reality without any hassle | 3.53 | 0.90 | 4 | 11 | 3 | .0207 | .207 |
| EE | My interaction with virtual reality is clear and understandable | 3.68 | 0.75 | 4 | 14 | 2 | .0009 | .013 |
| EE | Learning the use of virtual reality is not difficult for me | 3.84 | 0.69 | 4 | 15 | 1 | <.0001 | <.001 |
| FC | I can easily get technical support if I have problems using virtual reality | 3.58 | 0.84 | 4 | 11 | 2 | .0075 | .097 |
| FC | I know whom to contact if I experience any problems in using virtual reality | 3.05 | 1.03 | 3 | 7 | 6 | .8256 | 1.000 |
| FC | If I have any problems using virtual reality, I can reach the necessary information for a solution | 3.68 | 0.82 | 4 | 13 | 2 | .0019 | .026 |
4. Assumptions, reliability and what to distrust
Normality
A Shapiro-Wilk test rejects normality on 17 of the 18 items. That is expected rather than alarming: a five-point item measured on 19 people takes at most five distinct values, so it cannot look continuous. The question is whether it changes the answer. Wilcoxon signed-rank tests, which assume no distributional shape, agree with the t-tests on all four constructs.
| Construct | Shapiro-Wilk p | Non-normal | Wilcoxon p | t-test p |
|---|---|---|---|---|
| Performance Expectancy | .101 | no | .0341 | .0341 |
| Social Influence | .050 | yes | .6874 | .7215 |
| Effort Expectancy | .051 | no | .0006 | <.0001 |
| Facilitating Conditions | .301 | no | .0163 | .0151 |
Reliability
Cronbach's alpha is 0.912 across all 18 items, with a bootstrap interval of 0.833 to 0.949. Per construct it runs from 0.70 for Facilitating Conditions, which has only 3 items, to 0.93 for Social Influence.
Alpha is usually reported as a single number, and at this sample size that is misleading. Facilitating Conditions illustrates it: alpha is 0.70, which reads as acceptable, but its bootstrap interval runs from 0.047 to 0.882. An interval that spans from essentially no internal consistency to very high consistency is not evidence of anything, and the reason is that the construct has only 3 items measured on 19 people. Reported bare, 0.70 would imply a precision the data cannot support.
Alpha also rises with the number of items, so a three-item construct and a six-item construct are not directly comparable even setting precision aside. The whole-instrument figure of 0.912 is the most stable of them, and even it spans 0.833 to 0.949.
Distribution of totals
Totals range from 42 to 81 against a possible 18 to 90. The spread is the story: acceptance was not uniform, and a mean near the middle of that range describes a genuinely mixed group rather than a consensus.
5. Reproducing the published figures
The published article reports the four construct means and a UTAUT total. Recomputing them from the response file:
| Statistic | Published | Recomputed | Agrees |
|---|---|---|---|
| Performance Expectancy | 3.4298 | 3.43 | yes |
| Social Influence | 3.0788 | 3.08 | yes |
| Effort Expectancy | 3.7368 | 3.74 | yes |
| Facilitating Conditions | 3.4385 | 3.44 | yes |
| UTAUT total, SD | 10.6 | 10.60 | yes |
| UTAUT total, range | 42 to 81 | 42 to 81 | yes |
| UTAUT total, mean | 62.1 | 61.89 | no |
All four construct means reproduce to four decimal places, as do the total's standard deviation and range. The total mean does not: the response file gives 61.89 where the article reports 62.1.
The pattern points to a transcription slip rather than a different dataset. A different dataset would move the standard deviation and the construct means as well, and it moves neither. Reaching 62.1 would require the grand total across all 19 respondents to be about four points higher than the file contains.
This is recorded because a reader recomputing from the data will hit it. It does not affect the construct-level results, which are what the article's conclusions rest on.
6. Limitations
The design is a single group measured once, after the simulation. There is no pre-measurement and no comparison group, so nothing here describes change, only the level of acceptance afterwards. Testing against the scale midpoint asks whether students leaned positive, not whether the simulation moved them.
At 19 respondents from one nursing programme, precision is limited and generalisation beyond that programme is not supported. Acceptance is self-reported and was collected immediately after the session, so it describes an impression at that moment rather than sustained use. The instrument measures intention and perception, which are not the same as learning.
7. How this was computed
Everything above is computed in Python from the survey response file, with pandas and SciPy. Bootstrap intervals use 10,000 resamples with the seed fixed at 42; changing the seed changes the intervals in the last decimal. Where a bootstrap resample of a five-point item came out with every value identical, the standardised effect is undefined for that resample and it was dropped rather than allowed to distort the interval. Fewer than ten of 10,000 resamples were affected on any item.
The code is on the code page and in the repository, which includes the full rebuild pipeline.