Statistical analysis

Acceptance of AI-Enabled Virtual Reality

Summary

19 baccalaureate nursing students rated AI-generated virtual reality simulations on the 18-item UTAUT instrument, each item scored 1 to 5. Each construct was tested against the scale midpoint of 3 with a one-sample t-test.

Effort Expectancy is well above neutral (3.74, 95% CI 3.45 to 4.03, d = 1.23, p = <.0001). Performance Expectancy (3.43, p = .0341) and Facilitating Conditions (3.44, p = .0151) are modestly above it. Social Influence is not distinguishable from neutral (3.08, 95% CI 2.62 to 3.54, p = .7215).

Internal consistency is high for the instrument as a whole (alpha 0.912, 95% CI 0.833 to 0.949). At 19 respondents every interval here is wide, and that is the main thing to carry away.

1. The sample and the data

19 respondents answered all 18 items, with no missing values. Every response is an integer from 1 to 5. The response file carries no participant identifier, no timestamp and no free text, so nothing about it is disclosive on its own.

Two things follow from a sample of 19. Confidence intervals are wide, so small differences between constructs are not distinguishable. And because 18 items are scored on five points, the response distributions are lumpy, which matters for the normality assumption in section 4.

2. Acceptance by construct

The comparison of interest is against the scale midpoint of 3, where a respondent neither agrees nor disagrees. A construct score is the mean of its items, so it stays on the original 1 to 5 scale. Cohen's d is the distance from the midpoint in standard deviations, and its interval is a percentile bootstrap over 10,000 resamples.

Acceptance by UTAUT construct Mean score with 95% confidence interval. The dashed line is the scale midpoint (3). 2.5 3 neutral 3.5 4 Performance Expectancy 3.43 Social Influence 3.08 Effort Expectancy 3.74 Facilitating Conditions 3.44 Mean response (1 to 5 scale)
Figure 1. Mean score per construct against the scale midpoint.
One-sample comparisons against the midpoint of 3. Alpha is Cronbach's alpha for that construct's items.
ConstructItemsk Mean95% CId t (df)pα
PE Performance ExpectancyQ1-Q663.433.04 to 3.820.532.29 (18).03410.91
SI Social InfluenceQ7-Q1043.082.62 to 3.540.080.36 (18).72150.93
EE Effort ExpectancyQ11-Q1553.743.45 to 4.031.235.35 (18)<.00010.84
FC Facilitating ConditionsQ16-Q1833.443.10 to 3.780.622.69 (18).01510.70

The pattern is more interesting than a single acceptance score would be. Students found the technology easy to use, and that is the strongest finding by some distance. They were less convinced it would improve their performance, and they did not report social encouragement to use it at all. For a technology being introduced into a curriculum, those are different problems with different remedies.

3. Every item

Item-level results are exploratory. Eighteen tests on eighteen items inflates the familywise error rate, so Holm-adjusted p-values are given alongside the unadjusted ones. 5 of 18 items survive the adjustment.

Every item, grouped by construct Mean with 95% confidence interval. Rules separate the four constructs. 2 3 neutral 4 PE Using virtual reality increases my chances o… 3.63 Using virtual reality allows me to take resp… 3.63 Using virtual reality makes my life easier 3.47 Using virtual reality increases my productiv… 3.47 I find virtual reality useful for my daily l… 3.32 Using virtual reality enables me to accompli… 3.05 SI Most people who are important to me find it … 3.32 Most people who are important to me think I … 3.21 Most people who are important to me encourag… 2.89 Most people who are important to me use virt… 2.89 EE Learning how to use Virtual Reality is easy … 3.84 Learning the use of virtual reality is not d… 3.84 I find virtual reality easy to use 3.79 My interaction with virtual reality is clear… 3.68 I can use virtual reality without any hassle 3.53 FC If I have any problems using virtual reality… 3.68 I can easily get technical support if I have… 3.58 I know whom to contact if I experience any p… 3.05 Mean response (1 to 5 scale)
Figure 2. Every item with its 95% confidence interval, grouped by construct.
Where students agreed, and where they did not Share disagreeing (1 or 2) to the left, agreeing (4 or 5) to the right. n = 19. Disagree Agree PE Using virtual reality increases my chances o… 14/19 PE Using virtual reality makes my life easier 11/19 PE Using virtual reality increases my productiv… 11/19 PE Using virtual reality allows me to take resp… 10/19 PE I find virtual reality useful for my daily l… 11/19 PE Using virtual reality enables me to accompli… 8/19 SI Most people who are important to me find it … 9/19 SI Most people who are important to me think I … 7/19 SI Most people who are important to me encourag… 6/19 SI Most people who are important to me use virt… 5/19 EE Learning how to use Virtual Reality is easy … 15/19 EE Learning the use of virtual reality is not d… 15/19 EE I find virtual reality easy to use 15/19 EE My interaction with virtual reality is clear… 14/19 EE I can use virtual reality without any hassle 11/19 FC If I have any problems using virtual reality… 13/19 FC I can easily get technical support if I have… 11/19 FC I know whom to contact if I experience any p… 7/19 neutral
Figure 3. Counts agreeing (4 or 5) against disagreeing (1 or 2). Neutral responses are omitted, so the bars do not sum to 19.
Item-level descriptives and one-sample tests against the midpoint. Agree counts responses of 4 or 5; disagree counts 1 or 2.
Con.ItemMeanSD MdnAgreeDisagree pp Holm
PEUsing virtual reality increases my chances of solving problems I come across3.630.964144.0099.119
PEUsing virtual reality makes my life easier3.470.904112.0349.314
PEUsing virtual reality enables me to accomplish tasks more quickly3.050.97386.81581.000
PEUsing virtual reality increases my productivity3.470.964114.0462.370
PEI find virtual reality useful for my daily life3.321.064115.20921.000
PEUsing virtual reality allows me to take responsibility for my own learning3.631.074103.0187.206
SIMost people who are important to me encourage me to use virtual reality for learning purposes2.891.15369.69451.000
SIMost people who are important to me use virtual reality for learning purposes2.890.94358.63011.000
SIMost people who are important to me think I should use virtual reality for learning purposes3.211.18377.44771.000
SIMost people who are important to me find it helpful to use virtual reality for learning purposes3.320.89394.1374.962
EELearning how to use Virtual Reality is easy for me3.840.694151<.0001<.001
EEI find virtual reality easy to use3.790.794152.0004.006
EEI can use virtual reality without any hassle3.530.904113.0207.207
EEMy interaction with virtual reality is clear and understandable3.680.754142.0009.013
EELearning the use of virtual reality is not difficult for me3.840.694151<.0001<.001
FCI can easily get technical support if I have problems using virtual reality3.580.844112.0075.097
FCI know whom to contact if I experience any problems in using virtual reality3.051.03376.82561.000
FCIf I have any problems using virtual reality, I can reach the necessary information for a solution3.680.824132.0019.026

4. Assumptions, reliability and what to distrust

Normality

A Shapiro-Wilk test rejects normality on 17 of the 18 items. That is expected rather than alarming: a five-point item measured on 19 people takes at most five distinct values, so it cannot look continuous. The question is whether it changes the answer. Wilcoxon signed-rank tests, which assume no distributional shape, agree with the t-tests on all four constructs.

Normality of the construct scores, and a distribution-free comparison against the midpoint.
ConstructShapiro-Wilk pNon-normal Wilcoxon pt-test p
Performance Expectancy.101no.0341.0341
Social Influence.050yes.6874.7215
Effort Expectancy.051no.0006<.0001
Facilitating Conditions.301no.0163.0151

Reliability

Cronbach's alpha is 0.912 across all 18 items, with a bootstrap interval of 0.833 to 0.949. Per construct it runs from 0.70 for Facilitating Conditions, which has only 3 items, to 0.93 for Social Influence.

Alpha is usually reported as a single number, and at this sample size that is misleading. Facilitating Conditions illustrates it: alpha is 0.70, which reads as acceptable, but its bootstrap interval runs from 0.047 to 0.882. An interval that spans from essentially no internal consistency to very high consistency is not evidence of anything, and the reason is that the construct has only 3 items measured on 19 people. Reported bare, 0.70 would imply a precision the data cannot support.

Alpha also rises with the number of items, so a three-item construct and a six-item construct are not directly comparable even setting precision aside. The whole-instrument figure of 0.912 is the most stable of them, and even it spans 0.833 to 0.949.

Distribution of totals

Distribution of UTAUT total scores One value per respondent (n = 19). Possible range 18 to 90. 0 1 2 3 4 54 = all neutral mean 61.9 40 50 60 70 80 90 UTAUT total score
Figure 4. UTAUT total per respondent. A score of 54 would mean neutral on every item.

Totals range from 42 to 81 against a possible 18 to 90. The spread is the story: acceptance was not uniform, and a mean near the middle of that range describes a genuinely mixed group rather than a consensus.

5. Reproducing the published figures

The published article reports the four construct means and a UTAUT total. Recomputing them from the response file:

Published values against values recomputed from the survey responses.
StatisticPublishedRecomputedAgrees
Performance Expectancy3.42983.43yes
Social Influence3.07883.08yes
Effort Expectancy3.73683.74yes
Facilitating Conditions3.43853.44yes
UTAUT total, SD10.610.60yes
UTAUT total, range42 to 8142 to 81yes
UTAUT total, mean62.161.89no

All four construct means reproduce to four decimal places, as do the total's standard deviation and range. The total mean does not: the response file gives 61.89 where the article reports 62.1.

The pattern points to a transcription slip rather than a different dataset. A different dataset would move the standard deviation and the construct means as well, and it moves neither. Reaching 62.1 would require the grand total across all 19 respondents to be about four points higher than the file contains.

This is recorded because a reader recomputing from the data will hit it. It does not affect the construct-level results, which are what the article's conclusions rest on.

6. Limitations

The design is a single group measured once, after the simulation. There is no pre-measurement and no comparison group, so nothing here describes change, only the level of acceptance afterwards. Testing against the scale midpoint asks whether students leaned positive, not whether the simulation moved them.

At 19 respondents from one nursing programme, precision is limited and generalisation beyond that programme is not supported. Acceptance is self-reported and was collected immediately after the session, so it describes an impression at that moment rather than sustained use. The instrument measures intention and perception, which are not the same as learning.

7. How this was computed

Everything above is computed in Python from the survey response file, with pandas and SciPy. Bootstrap intervals use 10,000 resamples with the seed fixed at 42; changing the seed changes the intervals in the last decimal. Where a bootstrap resample of a five-point item came out with every value identical, the standardised effect is undefined for that resample and it was dropped rather than allowed to distort the interval. Fewer than ten of 10,000 resamples were affected on any item.

The code is on the code page and in the repository, which includes the full rebuild pipeline.