From Tool Adoption to Organisational Design
Research protocol for a survey on AI integration in agile software development
- Andreas Hinderks · Hannover University of Applied Sciences and Arts, Hannover, Germany
- Jörg Thomaschewski · University of Applied Sciences Emden/Leer, Emden, Germany
- Eva-Maria Schön · University of Applied Sciences Emden/Leer, Emden, Germany
Published as: Hinderks, A., Thomaschewski, J., & Schön, E.-M. (2026). From Tool Adoption to Organisational Design: An Empirical Study of AI Integration in Agile Software Development. ISD 2026, Prague.
Method at a glance
- Design
- Cross-sectional online survey
- Instrument
- 28 Likert items in four thematic blocks plus two open-text questions; 7-point scale normalised to −3 … +3
- Sample
- n = 99
- Population
- Professional software developers with agile experience and at least occasional use of AI coding tools
- Recruitment
- Prolific, screened convenience sample; three screening criteria, non-probabilistic opt-in
- Data collection
- March 2026
- Analysis
- Descriptive statistics per item: M, Mdn, SD, 95 % CI via the t-distribution
- Internal consistency per block: Cronbach's α, corrected item-total correlations, α if item deleted
- Exploratory factor analysis: PCA with Varimax rotation, Kaiser criterion
- Group comparisons: Welch's t-test or Mann-Whitney U depending on Shapiro-Wilk normality, Bonferroni-corrected
- Correlations: Spearman ρ, with Pearson r reported for reference
- Ethics and consent
- Participation was voluntary and compensated through Prolific. Responses carry the Prolific participant ID as a pseudonym for payment and quality control, but no direct identifiers such as name or e-mail; no individual participant is identifiable from anything reported here.
This is the full analysis behind our ISD 2026 paper. The paper reports the findings that fit within its page limit; this protocol documents the complete set — all 28 items with descriptive statistics, internal consistency for every block, the exploratory factor structure, and all four hypothesis tests, including the ones that did not work out.
It is supplementary material, not a paper. There is no discussion section and no argument to defend. Where the numbers are weak, they are shown as weak.
Info
Reading this alongside the paper. The paper paraphrases item wordings for space and points to an online appendix for the full text. This protocol is that appendix: the tables below use short forms so they stay readable, and the instrument section reproduces every item as it was put to respondents, along with the model description that preceded Block M.
Summary
- Perceived Changes
- +0.8495 % CI [+0.09, +1.59]
- How AI tools change coding workflows, roles, and team dynamics (C1–C8)
- Organizational Challenge
- +0.7395 % CI [+0.34, +1.12]
- Is AI integration a tool adoption or org design problem? (O1–O7)
- Model Plausibility
- +1.0995 % CI [+0.79, +1.39]
- Practitioner assessment of the proposed three-layer model (M1–M8)
- Agile in Transition
- +1.3195 % CI [+0.55, +2.07]
- Impact of AI tools on established agile practices (A1–A5)
Means across all items of a block, scale −3 to +3 (0 = neutral), n = 99.
Practitioners endorse the idea of AI development environments as shared infrastructure (O5: M = +1.38) and find the proposed three-layer model plausible in the abstract (M1: M = +1.49). They are markedly less confident that it would work where they actually work (M5: M = +0.76). That distance between plausible and applicable is the central observation of the study, and it is corroborated by M6: the same respondents agree that the model presupposes organisational maturity most companies do not yet have (M = +1.40).
Research questions
The survey addresses three questions, each operationalised by one or more measurement blocks:
- RQ1 — To what extent do practitioners perceive AI integration as an organisational design problem rather than a tool adoption problem? (Block O)
- RQ2 — How do they assess the three-layer model's plausibility? (Block M)
- RQ3 — What changes in roles, iterations, and agile practices are they experiencing? (Blocks C and A)
Four hypotheses were formulated in advance from the propositions of the position paper. They were not pre-registered; they are stated here as they were defined before analysis, and all four are reported below regardless of outcome.
Method
Sample
| Less than 1 year | 1 | |
|---|---|---|
| 1-2 years | 18 | |
| 3-5 years | 35 | |
| 6-10 years | 25 | |
| 11-15 years | 12 | |
| 16-20 years | 5 | |
| More than 20 years | 3 |
n = 99
| Software Developer / Engineer | 72 | |
|---|---|---|
| Software Architect | 2 | |
| Team Lead / Engineering Manager | 11 | |
| Scrum Master / Agile Coach | 0 | |
| Product Owner / Product Manager | 3 | |
| UX Designer / Researcher | 1 | |
| QA / Test Engineer | 1 | |
| DevOps / Platform Engineer | 2 | |
| Full-Stack Developer | 6 | |
| Other | 1 |
n = 99
| 1-3 | 21 | |
|---|---|---|
| 4-6 | 38 | |
| 7-10 | 29 | |
| 11-20 | 10 | |
| More than 20 | 1 |
n = 99
| Never | 0 | |
|---|---|---|
| Rarely | 2 | |
| Occasionally | 10 | |
| Frequently | 28 | |
| Daily | 59 |
n = 99
| 1-10 (Startup) | 18 | |
|---|---|---|
| 11-50 (Small) | 19 | |
| 51-200 (Medium) | 15 | |
| 201-1,000 (Large) | 16 | |
| More than 1,000 (Enterprise) | 31 |
n = 99
| GitHub Copilot | 62 | |
|---|---|---|
| Cursor | 28 | |
| Claude (Code / Chat) | 64 | |
| ChatGPT | 67 | |
| Gemini | 45 | |
| Windsurf / Codeium | 4 | |
| Amazon CodeWhisperer / Q | 5 | |
| Tabnine | 0 | |
| Other | 0 |
Multiple answers possible, 275 selections from 99 respondents
The sample is a screened convenience sample, not a representative draw from the developer population. Respondents had to be professional software developers, have experience with agile methods, and use AI-based coding tools at least occasionally. Two consequences follow directly from that screening and matter for every group comparison below: intensive AI users dominate (87 of 99 report daily or frequent use), and small teams dominate (88 of 99 work in teams of ten or fewer). Both grouping variables are therefore markedly imbalanced.
Instrument
The four blocks were defined a priori, each operationalising one proposition of the underlying position paper. This ensures content validity by construction, but it makes the blocks thematic item groups rather than reflective scales — a distinction that matters for how the reliability figures below should be read.
Block M is a special case: respondents did not rate the three-layer model as an abstraction. They were shown a concrete description of it — a shared "AI-Dev-System" maintained by System, Platform, and Product Teams — and rated the statements against that text. The wording of that stimulus therefore shapes every M result, and it is reproduced in full under the instrument.
Item wordings are shortened in the results tables below; hovering an item shows its full text, and the appendix lists all 28 verbatim.
Results
The scale runs from −3 (strongly disagree) to +3 (strongly agree), with 0 as the neutral midpoint. All items were answered by all 99 respondents.
Block C — Perceived changes from AI tools
| Item | Statement | M | Mdn | SD | 95 % CI | Distribution |
|---|---|---|---|---|---|---|
| C1 | AI tools have significantly changed how I approach coding tasks.Strong Agreement | +1.98 | +2.0 | 1.18 | [+1.74, +2.21] | |
| C2 | Implementation is no longer the main bottleneck in my work.Moderate Agreement | +0.94 | +1.0 | 1.56 | [+0.63, +1.25] | |
| C3 | My role has shifted from writing code to reviewing/steering AI outputs.Moderate Agreement | +0.72 | +1.0 | 1.69 | [+0.38, +1.05] | |
| C4 | AI tools allow my team to iterate much faster than before.Strong Agreement | +1.53 | +2.0 | 1.45 | [+1.24, +1.81] | |
| C5 | I spend more time on product decisions than on implementation.Moderate Agreement | +0.97 | +1.0 | 1.59 | [+0.65, +1.29] | |
| C6 | Different team members use different AI tools, which creates friction.Disagreement | −0.66 | −1.0 | 1.53 | [−0.96, −0.35] | |
| C7 | My team lacks clear guidelines on how to use AI tools effectively.Neutral | −0.22 | ±0.0 | 1.61 | [−0.54, +0.10] | |
| C8 | Pair programming with AI feels fundamentally different from human pairing.Moderate Agreement | +1.47 | +2.0 | 1.26 | [+1.22, +1.73] |
Two results stand out. C1 is the strongest agreement in the entire survey: AI tools have changed how respondents approach coding (M = +1.98). But C6 runs the other way — respondents disagree that different tools across team members create friction (M = −0.66), and C7 on missing guidelines sits at neutral (M = −0.22). The fragmentation the position paper anticipates is not yet experienced as a problem by the people in it.
| Item | Statement | Item-total r | α if deleted |
|---|---|---|---|
| C1 | AI tools have significantly changed how I approach coding tasks. | 0.665 | 0.709 |
| C2 | Implementation is no longer the main bottleneck in my work. | 0.683 | 0.693 |
| C3 | My role has shifted from writing code to reviewing/steering AI outputs. | 0.679 | 0.691 |
| C4 | AI tools allow my team to iterate much faster than before. | 0.684 | 0.696 |
| C5 | I spend more time on product decisions than on implementation. | 0.629 | 0.703 |
| C6* | Different team members use different AI tools, which creates friction. | 0.188 | 0.784 |
| C7* | My team lacks clear guidelines on how to use AI tools effectively. | 0.094 | 0.802 |
| C8* | Pair programming with AI feels fundamentally different from human pairing. | 0.183 | 0.777 |
Block O — AI integration as an organisational challenge
| Item | Statement | M | Mdn | SD | 95 % CI | Distribution |
|---|---|---|---|---|---|---|
| O1 | Integrating AI tools requires changes to team structures, not just tool adoption.Moderate Agreement | +0.70 | +1.0 | 1.57 | [+0.38, +1.01] | |
| O2 | Framing AI as "assistant"/"copilot" underestimates its organizational impact.Moderate Agreement | +0.74 | +1.0 | 1.60 | [+0.42, +1.06] | |
| O3 | Organization would benefit from dedicated teams for AI dev environments.Moderate Agreement | +1.13 | +1.0 | 1.33 | [+0.87, +1.40] | |
| O4 | Independent AI tool selection leads to fragmented tooling across the team.Neutral | +0.46 | +1.0 | 1.62 | [+0.14, +0.79] | |
| O5 | AI tools should be managed as shared infrastructure.Moderate Agreement | +1.38 | +2.0 | 1.27 | [+1.13, +1.64] | |
| O6 | Governance for AI-generated code is currently insufficient.Moderate Agreement | +0.62 | +1.0 | 1.54 | [+0.31, +0.92] | |
| O7 | AI integration is primarily a technical challenge, not organizational. [rev.]Neutral | +0.08 | ±0.0 | 1.57 | [−0.23, +0.39] |
Agreement is consistent but moderate. The strongest item is O5, managing AI tools as shared infrastructure (M = +1.38); the weakest is O7, the reverse-coded item on AI integration being primarily technical, which sits almost exactly at neutral (M = +0.08). Read together with C6, the picture is of practitioners who accept the organisational framing in principle without reporting the specific pain that motivates it.
| Item | Statement | Item-total r | α if deleted |
|---|---|---|---|
| O1 | Integrating AI tools requires changes to team structures, not just tool adoption. | 0.558 | 0.421 |
| O2 | Framing AI as "assistant"/"copilot" underestimates its organizational impact. | 0.414 | 0.483 |
| O3 | Organization would benefit from dedicated teams for AI dev environments. | 0.489 | 0.467 |
| O4* | Independent AI tool selection leads to fragmented tooling across the team. | 0.185 | 0.575 |
| O5 | AI tools should be managed as shared infrastructure. | 0.328 | 0.523 |
| O6* | Governance for AI-generated code is currently insufficient. | 0.075 | 0.611 |
| O7* | AI integration is primarily a technical challenge, not organizational. [rev.] | 0.095 | 0.606 |
Block M — Three-layer model plausibility
| Item | Statement | M | Mdn | SD | 95 % CI | Distribution |
|---|---|---|---|---|---|---|
| M1 | The three-layer model is a plausible way to organize AI-augmented dev.Moderate Agreement | +1.49 | +2.0 | 1.10 | [+1.28, +1.71] | |
| M2 | System Teams would reduce the 'paradox of choice' problem.Moderate Agreement | +1.39 | +2.0 | 1.22 | [+1.15, +1.64] | |
| M3 | Platform Teams would help my team work more consistently.Moderate Agreement | +1.27 | +2.0 | 1.33 | [+1.01, +1.54] | |
| M4 | I would prefer consuming a curated AI environment over selecting tools myself.Moderate Agreement | +0.80 | +1.0 | 1.70 | [+0.46, +1.14] | |
| M5 | This model would work in my current organizational context.Moderate Agreement | +0.76 | +1.0 | 1.53 | [+0.45, +1.06] | |
| M6 | This model requires organizational maturity most companies don't have yet.Moderate Agreement | +1.40 | +2.0 | 1.38 | [+1.13, +1.68] | |
| M7 | Small teams could implement a simplified version of this model.Moderate Agreement | +1.12 | +1.0 | 1.55 | [+0.81, +1.43] | |
| M8 | This model risks creating new silos between the three team types.Moderate Agreement | +0.52 | +1.0 | 1.38 | [+0.24, +0.79] |
M1 and M5 carry the central finding. The model is judged plausible (M = +1.49) far more readily than deployable in the respondent's own organisation (M = +0.76). M6 explains the gap in the respondents' own terms: the model requires a maturity most organisations lack (M = +1.40).
| Item | Statement | Item-total r | α if deleted |
|---|---|---|---|
| M1 | The three-layer model is a plausible way to organize AI-augmented dev. | 0.642 | 0.464 |
| M2 | System Teams would reduce the 'paradox of choice' problem. | 0.645 | 0.451 |
| M3 | Platform Teams would help my team work more consistently. | 0.661 | 0.433 |
| M4 | I would prefer consuming a curated AI environment over selecting tools myself. | 0.465 | 0.484 |
| M5 | This model would work in my current organizational context. | 0.496 | 0.478 |
| M6* | This model requires organizational maturity most companies don't have yet. | -0.033 | 0.643 |
| M7 | Small teams could implement a simplified version of this model. | 0.202 | 0.582 |
| M8* | This model risks creating new silos between the three team types. | -0.414 | 0.734 |
Block A — Agile practices in transition
| Item | Statement | M | Mdn | SD | 95 % CI | Distribution |
|---|---|---|---|---|---|---|
| A1 | The two-week sprint is becoming less meaningful.Neutral | +0.33 | +1.0 | 1.77 | [−0.02, +0.69] | |
| A2 | TDD is becoming more important as specification for AI-generated code.Strong Agreement | +1.61 | +2.0 | 1.14 | [+1.38, +1.83] | |
| A3 | Retrospectives should evaluate the AI dev environment, not just processes.Moderate Agreement | +1.29 | +1.0 | 1.35 | [+1.02, +1.56] | |
| A4 | CI is more critical because AI code requires stronger quality checks.Strong Agreement | +1.98 | +2.0 | 1.07 | [+1.77, +2.19] | |
| A5 | Most important skill is shifting from coding to evaluating/directing AI.Moderate Agreement | +1.33 | +1.0 | 1.47 | [+1.04, +1.63] |
The change is selective, not general. Continuous integration (A4: M = +1.98) and test-driven development (A2: M = +1.61) gain importance as verification mechanisms for AI-generated code. The two-week sprint, by contrast, is the only item in the block that does not clear the neutral band (A1: M = +0.33, CI including zero) — its confidence interval does not exclude "no change at all".
| Item | Statement | Item-total r | α if deleted |
|---|---|---|---|
| A1 | The two-week sprint is becoming less meaningful. | 0.424 | 0.512 |
| A2 | TDD is becoming more important as specification for AI-generated code. | 0.313 | 0.569 |
| A3 | Retrospectives should evaluate the AI dev environment, not just processes. | 0.393 | 0.527 |
| A4 | CI is more critical because AI code requires stronger quality checks. | 0.279 | 0.583 |
| A5 | Most important skill is shifting from coding to evaluating/directing AI. | 0.393 | 0.525 |
All items at a glance
Internal consistency — how to read these figures
Two of the four blocks fall below the conventional α ≥ .70 threshold, and one sits at .601. Reported as scale reliabilities, these would be poor results. They are not scale reliabilities.
The blocks were built from the propositions being tested, not extracted from the data. Block O deliberately mixes the organisational framing (O1, O2, O3) with items about currently perceived deficits (O6, O7); Block M deliberately pairs endorsement of the model (M1–M5) with two counter-pole items (M6 on required maturity, M8 on the risk of new silos). Those counter-pole items produce the negative item-total correlations visible in the tables — M8 at −0.414 is the clearest case. That is the intended structure of the instrument showing up in the statistics, not measurement error.
The practical consequence: interpretation proceeds at the item and sub-cluster level throughout this protocol. Block means are reported as summaries of what respondents said across a theme, not as scores on a latent construct.
Factor structure
Sampling adequacy: KMO = 0.757 (middling) · Bartlett’s test of sphericity: χ² = 1165.6, p < .001
The Kaiser criterion (eigenvalue > 1) retains 8 factors – considerably more than the 4 thematic blocks of the questionnaire. The matrix shows the first 4 factors after Varimax rotation.
| Factor | Eigenvalue | % variance | Cumulative |
|---|---|---|---|
| 1 | 6.466 | 23.1 % | 23.1 % |
| 2 | 3.446 | 12.3 % | 35.4 % |
| 3 | 2.153 | 7.7 % | 43.1 % |
| 4 | 1.920 | 6.9 % | 49.9 % |
| 5 | 1.409 | 5.0 % | 55.0 % |
| 6 | 1.310 | 4.7 % | 59.7 % |
| 7 | 1.213 | 4.3 % | 64.0 % |
| 8← Kaiser | 1.017 | 3.6 % | 67.6 % |
| 9 | 0.894 | 3.2 % | 70.8 % |
| 10 | 0.804 | 2.9 % | 73.7 % |
| 11 | 0.754 | 2.7 % | 76.4 % |
| 12 | 0.741 | 2.6 % | 79.0 % |
| 13 | 0.664 | 2.4 % | 81.4 % |
| 14 | 0.592 | 2.1 % | 83.5 % |
| 15 | 0.551 | 2.0 % | 85.5 % |
| 16 | 0.530 | 1.9 % | 87.4 % |
| 17 | 0.477 | 1.7 % | 89.1 % |
| 18 | 0.441 | 1.6 % | 90.7 % |
| 19 | 0.414 | 1.5 % | 92.1 % |
| 20 | 0.394 | 1.4 % | 93.5 % |
| 21 | 0.305 | 1.1 % | 94.6 % |
| Item | Factor 1 | Factor 2 | Factor 3 | Factor 4 | Primary |
|---|---|---|---|---|---|
| C1 | −0.819 | −0.090 | +0.033 | −0.106 | 1 |
| C2 | −0.822 | −0.060 | −0.017 | −0.041 | 1 |
| C3 | −0.798 | +0.030 | +0.115 | +0.002 | 1 |
| C4 | −0.857 | −0.071 | −0.001 | −0.050 | 1 |
| C5 | −0.836 | −0.008 | +0.047 | +0.018 | 1 |
| C6 | −0.147 | +0.039 | +0.603 | +0.111 | 3 |
| C7 | +0.040 | −0.029 | +0.557 | −0.354 | 3 |
| C8 | −0.133 | −0.004 | +0.106 | −0.475 | 4 |
| O1 | −0.315 | −0.056 | +0.644 | −0.099 | 3 |
| O2 | −0.515 | +0.036 | +0.478 | +0.062 | 1 |
| O3 | −0.523 | −0.215 | +0.295 | +0.056 | 1 |
| O4 | +0.092 | −0.145 | +0.435 | +0.102 | 3 |
| O5 | −0.193 | −0.268 | +0.329 | +0.060 | 3 |
| O6 | −0.149 | −0.111 | +0.102 | −0.352 | 4 |
| O7 | −0.119 | −0.024 | +0.187 | −0.027 | 3 |
| M1 | −0.183 | −0.527 | −0.030 | −0.024 | 2 |
| M2 | −0.128 | −0.522 | +0.010 | −0.166 | 2 |
| M3 | −0.105 | −0.494 | +0.164 | −0.025 | 2 |
| M4 | −0.007 | −0.414 | +0.137 | +0.118 | 2 |
| M5 | −0.065 | −0.421 | +0.073 | +0.038 | 2 |
| M6 | −0.157 | −0.081 | −0.012 | −0.401 | 4 |
| M7 | −0.214 | −0.232 | −0.137 | +0.130 | 2 |
| M8 | −0.063 | +0.253 | +0.165 | −0.347 | 4 |
| A1 | −0.460 | −0.050 | +0.155 | +0.074 | 1 |
| A2 | −0.364 | −0.168 | −0.274 | −0.179 | 1 |
| A3 | −0.315 | −0.351 | −0.182 | −0.052 | 2 |
| A4 | −0.340 | −0.222 | −0.117 | −0.375 | 4 |
| A5 | −0.694 | −0.105 | +0.205 | −0.069 | 1 |
The analysis confirms the reading above. The Kaiser criterion retains eight factors rather than the four thematic blocks, and the first four explain just under half the variance. Block C splits cleanly: C1–C5 load strongly on Factor 1, while C6 and C7 — the fragmentation and governance items — move to Factor 3 together with O1 and O4. The counter-pole items M6 and M8 leave Factor 2 and group on Factor 4.
In other words, the data separate experienced change from perceived organisational deficit, cutting across the block boundaries the questionnaire imposed. Any future version of this instrument should treat these as distinct sub-scales and validate them separately.
Hypothesis tests
H1 AI Usage Frequency → Organizational Framing (Block O)
Practitioners who use AI tools more frequently perceive AI integration as an organizational (not just technical) challenge.
| Group | n | M | SD | 95 % CI |
|---|---|---|---|---|
| HIGH (Daily + Frequently) | 87 | +0.84 | 0.65 | [+0.70, +0.98] |
| LOW (Occasionally + Rarely) | 12 | −0.25 | 0.83 | [−0.78, +0.28] |
- Normality
- HIGH p = 0.131, LOW p = 0.806 → both normal → parametric
- Test
- Welch's t-test: t(12.9) = 4.361, p < .001
- Effect size
- d = 1.625 (large)
Limited power: LOW n < 15 — limited power. Read the effect size with that in mind.
H2 Team Size → Model Plausibility (Block M)
Practitioners in larger teams rate the three-layer model as more plausible than those in smaller teams.
| Group | n | M | SD | 95 % CI |
|---|---|---|---|---|
| SMALL (1–3, 4–6, 7–10) | 88 | +1.12 | 0.68 | [+0.97, +1.26] |
| LARGE (11–20, >20) | 11 | +0.90 | 0.96 | [+0.25, +1.54] |
- Normality
- SMALL p = 0.013, LARGE p = 0.461 → non-normal → non-parametric
- Test
- Mann-Whitney U: U = 537.5, p = 0.554
- Effect size
- r = -0.111 (small)
Limited power: LARGE n < 15 — limited power. Read the effect size with that in mind.
Show item-level resultsHide item-level results
| Item | n S | M S | n L | M L | Test | Statistic | p | Effect | Verdict |
|---|---|---|---|---|---|---|---|---|---|
| M1 | 88 | +1.51 | 11 | +1.36 | Mann-Whitney U | U = 436.5 | p = 0.575 | r = 0.098 (negligible) | NOT SUPPORTED |
| M2 | 88 | +1.39 | 11 | +1.46 | Mann-Whitney U | U = 441.5 | p = 0.624 | r = 0.088 (negligible) | NOT SUPPORTED |
| M3 | 88 | +1.32 | 11 | +0.91 | Mann-Whitney U | U = 495.0 | p = 0.903 | r = -0.023 (negligible) | NOT SUPPORTED |
| M5 | 88 | +0.74 | 11 | +0.91 | Mann-Whitney U | U = 472.5 | p = 0.900 | r = 0.024 (negligible) | NOT SUPPORTED |
H3 Tool Fragmentation (C6) ↔ Structural Change Need (O1, O3, O5)
Perceived tool fragmentation correlates positively with the perceived need for structural organizational changes.
- Valid pairs
- n = 99
- Correlation
- Spearman ρ = 0.381, p < .001, 95 % CI [0.198, 0.538]Pearson r = 0.358 for reference
- Normality
- C6 p = 0.000, O_struct p = 0.000 → non-normal → Spearman ρ used
- Effect size
- moderate (Cohen, 1988)
H4 AI Usage Frequency → Role Shift (C3, C5)
Frequent AI users report a stronger shift from coder to reviewer/steward (C3) and from implementation to product decisions (C5). Bonferroni-adjusted α = 0.025.
Show item-level resultsHide item-level results
| Item | n H | M HIGH ± SD | CI H | n L | M LOW ± SD | CI L | Test | Statistic | p | Effect | Verdict |
|---|---|---|---|---|---|---|---|---|---|---|---|
| C3 | 87 | +1.02 ± 1.52 | [+0.70, +1.35] | 12 | -1.50 ± 1.09 | [-2.19, -0.81] | Mann-Whitney U | U = 934.5 | p < .001 | r = -0.790 (large) | SUPPORTED |
| C5 | 87 | +1.24 ± 1.40 | [+0.94, +1.54] | 12 | -1.00 ± 1.54 | [-1.98, -0.02] | Mann-Whitney U | U = 889.5 | p < .001 | r = -0.704 (large) | SUPPORTED |
H1 and H4 produce large effects, but both compare 87 respondents against 12, and that imbalance is a direct consequence of the screening criteria. The effects are reported because they were tested, not because the design supports strong claims about their magnitude. H2 found no team-size effect at all — neither on the block mean nor on any individual item.
Open-text responses
What would need to change in your organization?
67 responses · 30 words on average
| Team & Collaboration | 66 | |
|---|---|---|
| Organization & Culture | 47 | |
| Training & Skills | 25 | |
| AI Tools & Quality | 23 | |
| Governance & Guidelines | 12 |
Frequent terms
- teams 27
- team 19
- system 17
- organization 15
- model 13
- work 13
- tools 12
- shared 10
- platform 10
- training 10
- people 10
- change 9
- company 9
- working 9
- structure 9
Frequent phrases
- dedicated teams 4
- platform teams 4
- model work 3
- individual tool 3
- prompt engineering 3
- product teams 3
- team structure 3
- shared knowledge 2
Selected responses
„transitioning to this model requires shifting from individual tool adoption to a centralized platform engineering mindset where AI is managed as shared infrastructure rather than a personal assistant.“
„We would need to employ more people, our operations team is understaffed as it is, plus we don't have people experienced enough to design, implement and maintain such a system.“
„I think AI progress would have to slow down before we consider implementing this. With the current pace of innovation, it's better to let everyone have their own setup, at least for small startups like the one I work in.“
„I think that even two levels would be enough. In my organization, we are actively experimenting with a similar approach, and I really like it. The main challenge for organization is the budget (to have a separate team for creating and maintaining AI framework)“
„Integrating all members into a single system with shared AI tools would be the biggest challenge, since, as the company is small, it relies on each developer having the freedom to choose their tools, which would take some time to change the mindset“
„We would have to go slower to implement this - we don't have time to wait for dependencies between these teams“
6 of 67 responses, selected to show the range of positions – supportive as well as sceptical. Reproduced unchanged, including typos. The full set of responses is not published.
Which agile practice has changed the most?
58 responses · 18 words on average
| AI Tools & Quality | 40 | |
|---|---|---|
| Process & Agile | 28 | |
| Testing & CI/CD | 15 | |
| Team & Collaboration | 7 | |
| Training & Skills | 6 |
Frequent terms
- code 22
- changed 14
- time 11
- tools 10
- development 9
- work 9
- faster 8
- now 8
- sprint 7
- developers 6
- agile 6
- planning 6
- since 6
- reviews 5
- rather 5
Frequent phrases
- spend time 4
- code review 4
- code reviews 3
- nothing changed 2
- changed tools 2
- generate code 2
- reviewing validating 2
- refining ai-generated 2
Selected responses
„Code reviews have changed the most. I now spend less time on syntax and more time auditing AI logic for accuracy, security, and architectural fit.“
„I feel now the need for test driven development has increased a lot since AI has become a permanent part of a developer's life, since the code generated by AI needs a lot of evaluation and testing before it can be pushed to the production site.“
„The two-week sprint has become less meaningful due to the integration of AI.“
„Personally the team's agile practice remained the same even though all developers and PO use AI. It's more of AI helping us in bouncing off ideas and developers use to quickly create a quick POC for our clients. But agile methods remain the same for us“
4 of 58 responses, selected to show the range of positions – supportive as well as sceptical. Reproduced unchanged, including typos. The full set of responses is not published.
Limitations
Imbalanced comparison groups. The screening criteria produced 87 intensive against 12 occasional AI users, and 88 small against 11 large teams. H1, H2 and H4 all rest on these splits. The effect sizes should be read as indicative, not as precise estimates.
Blocks are not scales. As set out above, the four blocks are thematic groupings. Block means summarise a theme; they do not measure a latent construct, and they should not be treated as scale scores in secondary analysis.
Self-report on a contested topic. All measures are perceptions. Whether respondents' roles have actually shifted from producing to reviewing code is not observed here — only that they report it.
Convenience sample. Prolific respondents who opt into a survey about AI tools are unlikely to be representative of developers generally, and the sample skews towards smaller teams and fewer years of experience.
Single measurement point. The data are cross-sectional and were collected in March 2026, in a fast-moving field. Several respondents noted in the open text that their answer would depend on how quickly the tools keep changing.
The instrument
The complete questionnaire as respondents saw it, exported from LimeSurvey. This includes the response options nobody selected — they were on offer, and that they stayed empty is itself a finding: no respondent identified as a Scrum Master or Agile Coach, none reported never using AI tools (the screening excluded them), and Tabnine went unused.
Two questions were asked but not analysed in this study: which agile frameworks the team uses (A5) and whether the team works with agile or lean methods at all (A6). Both served as screening and context.
AI-Dev-Systems: Organizational Design for AI-Augmented Software Development
Welcome!
Thank you for participating in this research study.
We are investigating how AI tools are changing software development practices and team structures. The survey takes approximately 7 minutes to complete.
Your responses are anonymous and will be used for academic research only.
Researchers: Andreas Hinderks (Hannover University of Applied Sciences), Joerg Thomaschewski & Eva-Maria Schoen (Emden/Leer University of Applied Sciences)
Screening & Demographics
Please answer the following questions about your professional background and experience with AI tools.
Prolific Participant ID
Free-text field
How many years of professional experience do you have in software development?
- Less than 1 year
- 1-2 years
- 3-5 years
- 6-10 years
- 11-15 years
- 16-20 years
- More than 20 years
What is your current primary role?
- Software Developer / Engineer
- Software Architect
- Team Lead / Engineering Manager
- Scrum Master / Agile Coach
- Product Owner / Product Manager
- UX Designer / Researcher
- QA / Test Engineer
- DevOps / Platform Engineer
- Full-Stack Developer
- Other
How many people do you work with directly in your team?
- 1-3
- 4-6
- 7-10
- 11-20
- More than 20
Which agile framework(s) does your team use? (Select all that apply)
- Scrum
- Kanban
- Extreme Programming (XP)
- SAFe
- LeSS
- Nexus
- Shape Up
- Other
Do you currently work in a team that uses agile or lean methods?
- Yes
- No
Which AI-powered development tools do you currently use? (Select all that apply)
- GitHub Copilot
- Cursor
- Claude (Code / Chat)
- ChatGPT
- Gemini
- Windsurf / Codeium
- Amazon CodeWhisperer / Q
- Tabnine
- Other
How frequently do you use AI tools in your daily development work?
- Never
- Rarely
- Occasionally
- Frequently
- Daily
How large is your organization (total employees)?
- 1-10 (Startup)
- 11-50 (Small)
- 51-200 (Medium)
- 201-1,000 (Large)
- More than 1,000 (Enterprise)
Block 1: Perceived Changes from AI Tools
Please rate the following statements based on your personal experience with AI tools in software development.
Scale: 1 = Strongly Disagree ... 7 = Strongly Agree
- C1AI tools have significantly changed how I approach coding tasks.
- C2Since using AI tools, implementation is no longer the main bottleneck in my work.
- C3My role has shifted from writing code to reviewing and steering AI-generated outputs.
- C4AI tools allow my team to iterate much faster than before.
- C5I spend more time on product decisions (what to build) than on implementation (how to build it).
- C6Different team members use different AI tools, which creates friction.
- C7My team lacks clear guidelines on how to use AI tools effectively.
- C8Pair programming with AI feels fundamentally different from pair programming with a human colleague.
Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree
Block 2: AI Integration as an Organizational Challenge
The following statements address whether AI integration is primarily a tool adoption challenge or an organizational design problem.
Scale: 1 = Strongly Disagree ... 7 = Strongly Agree
- O1Integrating AI tools effectively requires changes to team structures, not just tool adoption.
- O2Framing AI as an "assistant" or "copilot" underestimates its organizational impact.
- O3Our organization would benefit from dedicated teams that design and maintain AI development environments.
- O4Each developer independently selecting their own AI tools leads to fragmented tooling across the team.
- O5AI tools should be managed as shared infrastructure rather than individual productivity tools.
- O6Governance (e.g., data privacy, quality standards) for AI-generated code is currently insufficient in my organization.
- O7Successful AI integration is primarily a technical challenge, not an organizational one. [reverse-coded]
Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree
Block 3: Three-Layer Model Plausibility
Please read the following description carefully, then rate the statements below.
Imagine that instead of each developer independently choosing and configuring their own AI tools, your organization creates a shared AI development environment — an "AI-Dev-System". This is a configured environment of AI agents, automated workflows, and curated knowledge bases that the entire organization uses for software development.
Three types of teams work together to make this system function:
System Teams design and build the AI-Dev-System. They select and configure AI agents, define interaction workflows, curate project-specific knowledge bases, and establish prompt engineering standards. Their goal is to encode organizational expertise into the system so that individual developers don't need deep AI configuration skills.
Platform Teams maintain and evolve the AI-Dev-System after deployment. They handle updates when foundation models change, fix issues that arise in daily use, extend capabilities based on team feedback, and manage regression testing. They ensure the system stays reliable and improves over time.
Product Teams develop software within the AI-Dev-System. rather than selecting AI tools individually, they consume curated capabilities provided by the platform — similar to how teams today use CI/CD pipelines without building them from scratch.
Scale: 1 = Strongly Disagree ... 7 = Strongly Agree
- M1The described three-layer model is a plausible way to organize AI-augmented software development.
- M2Having dedicated System Teams design AI environments would reduce the "paradox of choice" problem in my organization.
- M3Platform Teams maintaining AI-Dev-Systems would help my team work more consistently.
- M4As a Product Team member, I would prefer consuming a curated AI environment over selecting tools myself.
- M5This model would work in my current organizational context.
- M6This model requires a level of organizational maturity that most companies do not yet have.
- M7Small teams (fewer than 10 people) could implement a simplified version of this model.
- M8This model risks creating new silos between the three team types.
Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree
What would need to change in your organization to make a model like this work?
Free-text field · Please share your thoughts in a few sentences. Leave blank if you prefer not to answer.
Block 4: Agile Practices in Transition
Finally, please consider how AI tools are affecting established agile practices.
Scale: 1 = Strongly Disagree ... 7 = Strongly Agree
- A1The two-week sprint is becoming less meaningful as AI compresses implementation cycles.
- A2Test-driven development is becoming more important as a specification mechanism for AI-generated code.
- A3Retrospectives should evaluate the AI development environment, not just team processes.
- A4Continuous integration is more critical than before because AI-generated code requires stronger automated quality checks.
- A5The most important developer skill is shifting from coding to evaluating and directing AI outputs.
Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree
Which agile practice has changed the most since you started using AI tools, and how?
Free-text field · Please share your thoughts in a few sentences. Leave blank if you prefer not to answer.
Response scale
All statements were answered on the same seven-point scale. It was recoded symmetrically for the analysis:
| In the questionnaire | In the analysis |
|---|---|
| 1 – Strongly disagree | −3 |
| 2 – Disagree | −2 |
| 3 – Somewhat disagree | −1 |
| 4 – Neutral | ±0 |
| 5 – Somewhat agree | +1 |
| 6 – Agree | +2 |
| 7 – Strongly agree | +3 |
Methodology notes
- All Likert items use a 7-point scale, normalized to -3 to +3 (0 = Neutral)
- Original scale mapping: 1 → -3, 2 → -2, 3 → -1, 4 → 0, 5 → +1, 6 → +2, 7 → +3
- 95% confidence intervals computed using t-distribution (df = n-1)
- Item O7 is reverse-coded – positive values indicate disagreement with the reverse statement
- Distribution bars show percentage of responses per scale point
- Interpretation thresholds: Strong Agreement (M ≥ +1.5), Moderate Agreement (M ≥ +0.5), Neutral (M ≥ -0.5), Disagreement (M < -0.5)
- Cronbach's Alpha: internal consistency per block; α ≥ .70 acceptable, ≥ .80 good, ≥ .90 excellent
- Item-total correlation: Pearson r between item and sum of remaining block items (corrected)
- Exploratory Factor Analysis: PCA with Varimax rotation; Kaiser criterion for factor extraction
- KMO measure of sampling adequacy: ≥ .60 required, ≥ .80 meritorious
- Factor loadings ≥ |.40| considered salient for factor assignment
How to cite
In text
Hinderks, A., Thomaschewski, J. & Schön, E.-M. (2026). From Tool Adoption to Organisational Design: Research protocol for a survey on AI integration in agile software development [Research protocol, Version 1.0]. https://hinderks.org/studien/isd-2026-ai-integration
BibTeX
@misc{hinderks2026from,
author = {Andreas Hinderks and Jörg Thomaschewski and Eva-Maria Schön},
title = {{From Tool Adoption to Organisational Design: Research protocol for a survey on AI integration in agile software development}},
year = {2026},
howpublished = {Research protocol, Version 1.0},
url = {https://hinderks.org/studien/isd-2026-ai-integration},
note = {Accessed: <date>}
}- Version
- 1.0
- Licence
- CC BY 4.0
- Data
- Aggregated results in full on this page. Raw data are not published; verbatim open-text responses only as a curated selection.
Version history
- Version 1.0First publication of the protocol.
Frequently asked questions
The four blocks are thematic item groups, not reflective scales. They were defined in advance along the propositions being tested rather than extracted afterwards from a factor analysis. A low alpha here therefore indicates multidimensionality, not unreliable measurement. Blocks O and M also contain deliberate counter-pole items that correlate negatively with the rest of their block. Interpretation consequently proceeds at the item and sub-cluster level, not via block sums.
Respondents consider the three-layer model plausible in the abstract (M1: M = +1.49) but doubt it would work in their own organisation (M5: M = +0.76). This gap of 0.73 scale points between conceptual endorsement and perceived deployability is the central finding of the study. It is consistent with M6, where respondents largely agree that the model presupposes a level of organisational maturity most companies do not yet have.
Only to a limited degree. The effect d = 1.63 comes from a comparison of 87 intensive against 12 occasional AI users. The small comparison group limits statistical power and makes the effect estimate unstable, even though the difference itself is pronounced. The same caveat applies to H2 and H4. The imbalance does reflect the population, however: in a sample screened for developers who use AI tools, occasional users are naturally rare.
No. This protocol contains the aggregated results: descriptive statistics per item, reliability figures, factor loadings and test results. Person-level raw data are not published. Of the open-text responses, only a curated selection appears, chosen to reflect the range of positions.