Zum Inhalt springen
Supplementary materialVersion 1.0

From Tool Adoption to Organisational Design

Research protocol for a survey on AI integration in agile software development

  • Andreas Hinderks · Hannover University of Applied Sciences and Arts, Hannover, Germany
  • Jörg Thomaschewski · University of Applied Sciences Emden/Leer, Emden, Germany
  • Eva-Maria Schön · University of Applied Sciences Emden/Leer, Emden, Germany

Published as: Hinderks, A., Thomaschewski, J., & Schön, E.-M. (2026). From Tool Adoption to Organisational Design: An Empirical Study of AI Integration in Agile Software Development. ISD 2026, Prague.

Method at a glance

Design
Cross-sectional online survey
Instrument
28 Likert items in four thematic blocks plus two open-text questions; 7-point scale normalised to −3 … +3
Sample
n = 99
Population
Professional software developers with agile experience and at least occasional use of AI coding tools
Recruitment
Prolific, screened convenience sample; three screening criteria, non-probabilistic opt-in
Data collection
March 2026
Analysis
  • Descriptive statistics per item: M, Mdn, SD, 95 % CI via the t-distribution
  • Internal consistency per block: Cronbach's α, corrected item-total correlations, α if item deleted
  • Exploratory factor analysis: PCA with Varimax rotation, Kaiser criterion
  • Group comparisons: Welch's t-test or Mann-Whitney U depending on Shapiro-Wilk normality, Bonferroni-corrected
  • Correlations: Spearman ρ, with Pearson r reported for reference
Ethics and consent
Participation was voluntary and compensated through Prolific. Responses carry the Prolific participant ID as a pseudonym for payment and quality control, but no direct identifiers such as name or e-mail; no individual participant is identifiable from anything reported here.

This is the full analysis behind our ISD 2026 paper. The paper reports the findings that fit within its page limit; this protocol documents the complete set — all 28 items with descriptive statistics, internal consistency for every block, the exploratory factor structure, and all four hypothesis tests, including the ones that did not work out.

It is supplementary material, not a paper. There is no discussion section and no argument to defend. Where the numbers are weak, they are shown as weak.

Info

Reading this alongside the paper. The paper paraphrases item wordings for space and points to an online appendix for the full text. This protocol is that appendix: the tables below use short forms so they stay readable, and the instrument section reproduces every item as it was put to respondents, along with the model description that preceded Block M.

Summary

Perceived Changes
+0.8495 % CI [+0.09, +1.59]
How AI tools change coding workflows, roles, and team dynamics (C1–C8)
Organizational Challenge
+0.7395 % CI [+0.34, +1.12]
Is AI integration a tool adoption or org design problem? (O1–O7)
Model Plausibility
+1.0995 % CI [+0.79, +1.39]
Practitioner assessment of the proposed three-layer model (M1–M8)
Agile in Transition
+1.3195 % CI [+0.55, +2.07]
Impact of AI tools on established agile practices (A1–A5)

Means across all items of a block, scale −3 to +3 (0 = neutral), n = 99.

Practitioners endorse the idea of AI development environments as shared infrastructure (O5: M = +1.38) and find the proposed three-layer model plausible in the abstract (M1: M = +1.49). They are markedly less confident that it would work where they actually work (M5: M = +0.76). That distance between plausible and applicable is the central observation of the study, and it is corroborated by M6: the same respondents agree that the model presupposes organisational maturity most companies do not yet have (M = +1.40).

Research questions

The survey addresses three questions, each operationalised by one or more measurement blocks:

  • RQ1 — To what extent do practitioners perceive AI integration as an organisational design problem rather than a tool adoption problem? (Block O)
  • RQ2 — How do they assess the three-layer model's plausibility? (Block M)
  • RQ3 — What changes in roles, iterations, and agile practices are they experiencing? (Blocks C and A)

Four hypotheses were formulated in advance from the propositions of the position paper. They were not pre-registered; they are stated here as they were defined before analysis, and all four are reported below regardless of outcome.

Method

Sample

Professional experienceHow many years of professional experience do you have in software development?
Professional experience: How many years of professional experience do you have in software development?
Less than 1 year1
1-2 years18
3-5 years35
6-10 years25
11-15 years12
16-20 years5
More than 20 years3

n = 99

Primary roleWhat is your current primary role?
Primary role: What is your current primary role?
Software Developer / Engineer72
Software Architect2
Team Lead / Engineering Manager11
Scrum Master / Agile Coach0
Product Owner / Product Manager3
UX Designer / Researcher1
QA / Test Engineer1
DevOps / Platform Engineer2
Full-Stack Developer6
Other1

n = 99

Team sizeHow many people do you work with directly in your team?
Team size: How many people do you work with directly in your team?
1-321
4-638
7-1029
11-2010
More than 201

n = 99

AI usage frequencyHow frequently do you use AI tools in your daily development work?
AI usage frequency: How frequently do you use AI tools in your daily development work?
Never0
Rarely2
Occasionally10
Frequently28
Daily59

n = 99

Organisation sizeHow large is your organization (total employees)?
Organisation size: How large is your organization (total employees)?
1-10 (Startup)18
11-50 (Small)19
51-200 (Medium)15
201-1,000 (Large)16
More than 1,000 (Enterprise)31

n = 99

AI tools usedWhich AI-powered development tools do you currently use? (Select all that apply)
AI tools used: Which AI-powered development tools do you currently use? (Select all that apply)
GitHub Copilot62
Cursor28
Claude (Code / Chat)64
ChatGPT67
Gemini45
Windsurf / Codeium4
Amazon CodeWhisperer / Q5
Tabnine0
Other0

Multiple answers possible, 275 selections from 99 respondents

The sample is a screened convenience sample, not a representative draw from the developer population. Respondents had to be professional software developers, have experience with agile methods, and use AI-based coding tools at least occasionally. Two consequences follow directly from that screening and matter for every group comparison below: intensive AI users dominate (87 of 99 report daily or frequent use), and small teams dominate (88 of 99 work in teams of ten or fewer). Both grouping variables are therefore markedly imbalanced.

Instrument

The four blocks were defined a priori, each operationalising one proposition of the underlying position paper. This ensures content validity by construction, but it makes the blocks thematic item groups rather than reflective scales — a distinction that matters for how the reliability figures below should be read.

Block M is a special case: respondents did not rate the three-layer model as an abstraction. They were shown a concrete description of it — a shared "AI-Dev-System" maintained by System, Platform, and Product Teams — and rated the statements against that text. The wording of that stimulus therefore shapes every M result, and it is reproduced in full under the instrument.

Item wordings are shortened in the results tables below; hovering an item shows its full text, and the appendix lists all 28 verbatim.

Results

The scale runs from −3 (strongly disagree) to +3 (strongly agree), with 0 as the neutral midpoint. All items were answered by all 99 respondents.

Strongly disagree (-3)Disagree (-2)Somewhat disagree (-1)Neutral (0)Somewhat agree (+1)Agree (+2)Strongly agree (+3)

Block C — Perceived changes from AI tools

ItemStatementMMdnSD95 % CIDistribution
C1AI tools have significantly changed how I approach coding tasks.Strong Agreement+1.98+2.01.18[+1.74, +2.21]
C2Implementation is no longer the main bottleneck in my work.Moderate Agreement+0.94+1.01.56[+0.63, +1.25]
C3My role has shifted from writing code to reviewing/steering AI outputs.Moderate Agreement+0.72+1.01.69[+0.38, +1.05]
C4AI tools allow my team to iterate much faster than before.Strong Agreement+1.53+2.01.45[+1.24, +1.81]
C5I spend more time on product decisions than on implementation.Moderate Agreement+0.97+1.01.59[+0.65, +1.29]
C6Different team members use different AI tools, which creates friction.Disagreement−0.66−1.01.53[−0.96, −0.35]
C7My team lacks clear guidelines on how to use AI tools effectively.Neutral−0.22±0.01.61[−0.54, +0.10]
C8Pair programming with AI feels fundamentally different from human pairing.Moderate Agreement+1.47+2.01.26[+1.22, +1.73]
n = 99. Scale −3 (strongly disagree) to +3 (strongly agree), 0 = neutral. Confidence intervals via the t-distribution.

Two results stand out. C1 is the strongest agreement in the entire survey: AI tools have changed how respondents approach coding (M = +1.98). But C6 runs the other way — respondents disagree that different tools across team members create friction (M = −0.66), and C7 on missing guidelines sits at neutral (M = −0.22). The fragmentation the position paper anticipates is not yet experienced as a problem by the people in it.

Cronbach’s α = 0.76295 % CI [0.646, 0.830]Acceptable (α ≥ .70)8 items · n = 99
ItemStatementItem-total rα if deleted
C1AI tools have significantly changed how I approach coding tasks.0.6650.709
C2Implementation is no longer the main bottleneck in my work.0.6830.693
C3My role has shifted from writing code to reviewing/steering AI outputs.0.6790.691
C4AI tools allow my team to iterate much faster than before.0.6840.696
C5I spend more time on product decisions than on implementation.0.6290.703
C6*Different team members use different AI tools, which creates friction.0.1880.784
C7*My team lacks clear guidelines on how to use AI tools effectively.0.0940.802
C8*Pair programming with AI feels fundamentally different from human pairing.0.1830.777
Item-total r = corrected correlation of an item with the sum of the remaining items in its block. * marks 3 items with an item-total correlation below .20 (C6, C7, C8) – these measure a facet of their own rather than the same quantity as the rest of the block.

Block O — AI integration as an organisational challenge

ItemStatementMMdnSD95 % CIDistribution
O1Integrating AI tools requires changes to team structures, not just tool adoption.Moderate Agreement+0.70+1.01.57[+0.38, +1.01]
O2Framing AI as "assistant"/"copilot" underestimates its organizational impact.Moderate Agreement+0.74+1.01.60[+0.42, +1.06]
O3Organization would benefit from dedicated teams for AI dev environments.Moderate Agreement+1.13+1.01.33[+0.87, +1.40]
O4Independent AI tool selection leads to fragmented tooling across the team.Neutral+0.46+1.01.62[+0.14, +0.79]
O5AI tools should be managed as shared infrastructure.Moderate Agreement+1.38+2.01.27[+1.13, +1.64]
O6Governance for AI-generated code is currently insufficient.Moderate Agreement+0.62+1.01.54[+0.31, +0.92]
O7AI integration is primarily a technical challenge, not organizational. [rev.]Neutral+0.08±0.01.57[−0.23, +0.39]
n = 99. Scale −3 (strongly disagree) to +3 (strongly agree), 0 = neutral. Confidence intervals via the t-distribution.

Agreement is consistent but moderate. The strongest item is O5, managing AI tools as shared infrastructure (M = +1.38); the weakest is O7, the reverse-coded item on AI integration being primarily technical, which sits almost exactly at neutral (M = +0.08). Read together with C6, the picture is of practitioners who accept the organisational framing in principle without reporting the specific pain that motivates it.

Cronbach’s α = 0.57195 % CI [0.398, 0.676]Poor (α ≥ .50)7 items · n = 99
ItemStatementItem-total rα if deleted
O1Integrating AI tools requires changes to team structures, not just tool adoption.0.5580.421
O2Framing AI as "assistant"/"copilot" underestimates its organizational impact.0.4140.483
O3Organization would benefit from dedicated teams for AI dev environments.0.4890.467
O4*Independent AI tool selection leads to fragmented tooling across the team.0.1850.575
O5AI tools should be managed as shared infrastructure.0.3280.523
O6*Governance for AI-generated code is currently insufficient.0.0750.611
O7*AI integration is primarily a technical challenge, not organizational. [rev.]0.0950.606
Item-total r = corrected correlation of an item with the sum of the remaining items in its block. * marks 3 items with an item-total correlation below .20 (O4, O6, O7) – these measure a facet of their own rather than the same quantity as the rest of the block.

Block M — Three-layer model plausibility

ItemStatementMMdnSD95 % CIDistribution
M1The three-layer model is a plausible way to organize AI-augmented dev.Moderate Agreement+1.49+2.01.10[+1.28, +1.71]
M2System Teams would reduce the 'paradox of choice' problem.Moderate Agreement+1.39+2.01.22[+1.15, +1.64]
M3Platform Teams would help my team work more consistently.Moderate Agreement+1.27+2.01.33[+1.01, +1.54]
M4I would prefer consuming a curated AI environment over selecting tools myself.Moderate Agreement+0.80+1.01.70[+0.46, +1.14]
M5This model would work in my current organizational context.Moderate Agreement+0.76+1.01.53[+0.45, +1.06]
M6This model requires organizational maturity most companies don't have yet.Moderate Agreement+1.40+2.01.38[+1.13, +1.68]
M7Small teams could implement a simplified version of this model.Moderate Agreement+1.12+1.01.55[+0.81, +1.43]
M8This model risks creating new silos between the three team types.Moderate Agreement+0.52+1.01.38[+0.24, +0.79]
n = 99. Scale −3 (strongly disagree) to +3 (strongly agree), 0 = neutral. Confidence intervals via the t-distribution.

M1 and M5 carry the central finding. The model is judged plausible (M = +1.49) far more readily than deployable in the respondent's own organisation (M = +0.76). M6 explains the gap in the respondents' own terms: the model requires a maturity most organisations lack (M = +1.40).

Cronbach’s α = 0.58595 % CI [0.450, 0.674]Poor (α ≥ .50)8 items · n = 99
ItemStatementItem-total rα if deleted
M1The three-layer model is a plausible way to organize AI-augmented dev.0.6420.464
M2System Teams would reduce the 'paradox of choice' problem.0.6450.451
M3Platform Teams would help my team work more consistently.0.6610.433
M4I would prefer consuming a curated AI environment over selecting tools myself.0.4650.484
M5This model would work in my current organizational context.0.4960.478
M6*This model requires organizational maturity most companies don't have yet.-0.0330.643
M7Small teams could implement a simplified version of this model.0.2020.582
M8*This model risks creating new silos between the three team types.-0.4140.734
Item-total r = corrected correlation of an item with the sum of the remaining items in its block. * marks 2 items with an item-total correlation below .20 (M6, M8) – these measure a facet of their own rather than the same quantity as the rest of the block.

Block A — Agile practices in transition

ItemStatementMMdnSD95 % CIDistribution
A1The two-week sprint is becoming less meaningful.Neutral+0.33+1.01.77[−0.02, +0.69]
A2TDD is becoming more important as specification for AI-generated code.Strong Agreement+1.61+2.01.14[+1.38, +1.83]
A3Retrospectives should evaluate the AI dev environment, not just processes.Moderate Agreement+1.29+1.01.35[+1.02, +1.56]
A4CI is more critical because AI code requires stronger quality checks.Strong Agreement+1.98+2.01.07[+1.77, +2.19]
A5Most important skill is shifting from coding to evaluating/directing AI.Moderate Agreement+1.33+1.01.47[+1.04, +1.63]
n = 99. Scale −3 (strongly disagree) to +3 (strongly agree), 0 = neutral. Confidence intervals via the t-distribution.

The change is selective, not general. Continuous integration (A4: M = +1.98) and test-driven development (A2: M = +1.61) gain importance as verification mechanisms for AI-generated code. The two-week sprint, by contrast, is the only item in the block that does not clear the neutral band (A1: M = +0.33, CI including zero) — its confidence interval does not exclude "no change at all".

Cronbach’s α = 0.60195 % CI [0.432, 0.720]Questionable (α ≥ .60)5 items · n = 99
ItemStatementItem-total rα if deleted
A1The two-week sprint is becoming less meaningful.0.4240.512
A2TDD is becoming more important as specification for AI-generated code.0.3130.569
A3Retrospectives should evaluate the AI dev environment, not just processes.0.3930.527
A4CI is more critical because AI code requires stronger quality checks.0.2790.583
A5Most important skill is shifting from coding to evaluating/directing AI.0.3930.525
Item-total r = corrected correlation of an item with the sum of the remaining items in its block.

All items at a glance

Block C · Perceived Changes from AI ToolsC1C1: M = +1.98, 95 % CI [+1.74, +2.21] – AI tools have significantly changed how I approach coding tasks.C2C2: M = +0.94, 95 % CI [+0.63, +1.25] – Implementation is no longer the main bottleneck in my work.C3C3: M = +0.72, 95 % CI [+0.38, +1.05] – My role has shifted from writing code to reviewing/steering AI outputs.C4C4: M = +1.53, 95 % CI [+1.24, +1.81] – AI tools allow my team to iterate much faster than before.C5C5: M = +0.97, 95 % CI [+0.65, +1.29] – I spend more time on product decisions than on implementation.C6C6: M = −0.66, 95 % CI [−0.96, −0.35] – Different team members use different AI tools, which creates friction.C7C7: M = −0.22, 95 % CI [−0.54, +0.10] – My team lacks clear guidelines on how to use AI tools effectively.C8C8: M = +1.47, 95 % CI [+1.22, +1.73] – Pair programming with AI feels fundamentally different from human pairing.Block O · AI Integration as Organisational ChallengeO1O1: M = +0.70, 95 % CI [+0.38, +1.01] – Integrating AI tools requires changes to team structures, not just tool adoption.O2O2: M = +0.74, 95 % CI [+0.42, +1.06] – Framing AI as "assistant"/"copilot" underestimates its organizational impact.O3O3: M = +1.13, 95 % CI [+0.87, +1.40] – Organization would benefit from dedicated teams for AI dev environments.O4O4: M = +0.46, 95 % CI [+0.14, +0.79] – Independent AI tool selection leads to fragmented tooling across the team.O5O5: M = +1.38, 95 % CI [+1.13, +1.64] – AI tools should be managed as shared infrastructure.O6O6: M = +0.62, 95 % CI [+0.31, +0.92] – Governance for AI-generated code is currently insufficient.O7O7: M = +0.08, 95 % CI [−0.23, +0.39] – AI integration is primarily a technical challenge, not organizational. [rev.]Block M · Three-Layer Model PlausibilityM1M1: M = +1.49, 95 % CI [+1.28, +1.71] – The three-layer model is a plausible way to organize AI-augmented dev.M2M2: M = +1.39, 95 % CI [+1.15, +1.64] – System Teams would reduce the 'paradox of choice' problem.M3M3: M = +1.27, 95 % CI [+1.01, +1.54] – Platform Teams would help my team work more consistently.M4M4: M = +0.80, 95 % CI [+0.46, +1.14] – I would prefer consuming a curated AI environment over selecting tools myself.M5M5: M = +0.76, 95 % CI [+0.45, +1.06] – This model would work in my current organizational context.M6M6: M = +1.40, 95 % CI [+1.13, +1.68] – This model requires organizational maturity most companies don't have yet.M7M7: M = +1.12, 95 % CI [+0.81, +1.43] – Small teams could implement a simplified version of this model.M8M8: M = +0.52, 95 % CI [+0.24, +0.79] – This model risks creating new silos between the three team types.Block A · Agile Practices in TransitionA1A1: M = +0.33, 95 % CI [−0.02, +0.69] – The two-week sprint is becoming less meaningful.A2A2: M = +1.61, 95 % CI [+1.38, +1.83] – TDD is becoming more important as specification for AI-generated code.A3A3: M = +1.29, 95 % CI [+1.02, +1.56] – Retrospectives should evaluate the AI dev environment, not just processes.A4A4: M = +1.98, 95 % CI [+1.77, +2.19] – CI is more critical because AI code requires stronger quality checks.A5A5: M = +1.33, 95 % CI [+1.04, +1.63] – Most important skill is shifting from coding to evaluating/directing AI.-3-2-10+1+2+3strongly disagree ← neutral → strongly agree
Dot = mean, line = 95 % confidence interval. The individual figures are in the item tables of the respective blocks.

Internal consistency — how to read these figures

Two of the four blocks fall below the conventional α ≥ .70 threshold, and one sits at .601. Reported as scale reliabilities, these would be poor results. They are not scale reliabilities.

The blocks were built from the propositions being tested, not extracted from the data. Block O deliberately mixes the organisational framing (O1, O2, O3) with items about currently perceived deficits (O6, O7); Block M deliberately pairs endorsement of the model (M1–M5) with two counter-pole items (M6 on required maturity, M8 on the risk of new silos). Those counter-pole items produce the negative item-total correlations visible in the tables — M8 at −0.414 is the clearest case. That is the intended structure of the instrument showing up in the statistics, not measurement error.

The practical consequence: interpretation proceeds at the item and sub-cluster level throughout this protocol. Block means are reported as summaries of what respondents said across a theme, not as scores on a latent construct.

Factor structure

Sampling adequacy: KMO = 0.757 (middling) · Bartlett’s test of sphericity: χ² = 1165.6, p < .001

The Kaiser criterion (eigenvalue > 1) retains 8 factors – considerably more than the 4 thematic blocks of the questionnaire. The matrix shows the first 4 factors after Varimax rotation.

Eigenvalues and variance explained
FactorEigenvalue% varianceCumulative
16.46623.1 %23.1 %
23.44612.3 %35.4 %
32.1537.7 %43.1 %
41.9206.9 %49.9 %
51.4095.0 %55.0 %
61.3104.7 %59.7 %
71.2134.3 %64.0 %
8← Kaiser1.0173.6 %67.6 %
90.8943.2 %70.8 %
100.8042.9 %73.7 %
110.7542.7 %76.4 %
120.7412.6 %79.0 %
130.6642.4 %81.4 %
140.5922.1 %83.5 %
150.5512.0 %85.5 %
160.5301.9 %87.4 %
170.4771.7 %89.1 %
180.4411.6 %90.7 %
190.4141.5 %92.1 %
200.3941.4 %93.5 %
210.3051.1 %94.6 %
Rotated component matrix (Varimax)
ItemFactor 1Factor 2Factor 3Factor 4Primary
C10.8190.090+0.0330.1061
C20.8220.0600.0170.0411
C30.798+0.030+0.115+0.0021
C40.8570.0710.0010.0501
C50.8360.008+0.047+0.0181
C60.147+0.039+0.603+0.1113
C7+0.0400.029+0.5570.3543
C80.1330.004+0.1060.4754
O10.3150.056+0.6440.0993
O20.515+0.036+0.478+0.0621
O30.5230.215+0.295+0.0561
O4+0.0920.145+0.435+0.1023
O50.1930.268+0.329+0.0603
O60.1490.111+0.1020.3524
O70.1190.024+0.1870.0273
M10.1830.5270.0300.0242
M20.1280.522+0.0100.1662
M30.1050.494+0.1640.0252
M40.0070.414+0.137+0.1182
M50.0650.421+0.073+0.0382
M60.1570.0810.0120.4014
M70.2140.2320.137+0.1302
M80.063+0.253+0.1650.3474
A10.4600.050+0.155+0.0741
A20.3640.1680.2740.1791
A30.3150.3510.1820.0522
A40.3400.2220.1170.3754
A50.6940.105+0.2050.0691
Loadings of |λ| ≥ 0.40 are highlighted. The sign of a component is arbitrary in PCA and carries no substantive meaning; what counts is the magnitude.

The analysis confirms the reading above. The Kaiser criterion retains eight factors rather than the four thematic blocks, and the first four explain just under half the variance. Block C splits cleanly: C1–C5 load strongly on Factor 1, while C6 and C7 — the fragmentation and governance items — move to Factor 3 together with O1 and O4. The counter-pole items M6 and M8 leave Factor 2 and group on Factor 4.

In other words, the data separate experienced change from perceived organisational deficit, cutting across the block boundaries the questionnaire imposed. Any future version of this instrument should treat these as distinct sub-scales and validate them separately.

Hypothesis tests

H1 AI Usage Frequency → Organizational Framing (Block O)

Practitioners who use AI tools more frequently perceive AI integration as an organizational (not just technical) challenge.

Supported
GroupnMSD95 % CI
HIGH (Daily + Frequently)87+0.840.65[+0.70, +0.98]
LOW (Occasionally + Rarely)12−0.250.83[−0.78, +0.28]
Normality
HIGH p = 0.131, LOW p = 0.806 → both normal → parametric
Test
Welch's t-test: t(12.9) = 4.361, p < .001
Effect size
d = 1.625 (large)

Limited power: LOW n < 15 — limited power. Read the effect size with that in mind.

H2 Team Size → Model Plausibility (Block M)

Practitioners in larger teams rate the three-layer model as more plausible than those in smaller teams.

Not supported
GroupnMSD95 % CI
SMALL (1–3, 4–6, 7–10)88+1.120.68[+0.97, +1.26]
LARGE (11–20, >20)11+0.900.96[+0.25, +1.54]
Normality
SMALL p = 0.013, LARGE p = 0.461 → non-normal → non-parametric
Test
Mann-Whitney U: U = 537.5, p = 0.554
Effect size
r = -0.111 (small)

Limited power: LARGE n < 15 — limited power. Read the effect size with that in mind.

Show item-level results
Itemn SM Sn LM LTestStatisticpEffectVerdict
M188+1.5111+1.36Mann-Whitney UU = 436.5p = 0.575r = 0.098 (negligible)NOT SUPPORTED
M288+1.3911+1.46Mann-Whitney UU = 441.5p = 0.624r = 0.088 (negligible)NOT SUPPORTED
M388+1.3211+0.91Mann-Whitney UU = 495.0p = 0.903r = -0.023 (negligible)NOT SUPPORTED
M588+0.7411+0.91Mann-Whitney UU = 472.5p = 0.900r = 0.024 (negligible)NOT SUPPORTED

H3 Tool Fragmentation (C6) ↔ Structural Change Need (O1, O3, O5)

Perceived tool fragmentation correlates positively with the perceived need for structural organizational changes.

Supported
Valid pairs
n = 99
Correlation
Spearman ρ = 0.381, p < .001, 95 % CI [0.198, 0.538]Pearson r = 0.358 for reference
Normality
C6 p = 0.000, O_struct p = 0.000 → non-normal → Spearman ρ used
Effect size
moderate (Cohen, 1988)

H4 AI Usage Frequency → Role Shift (C3, C5)

Frequent AI users report a stronger shift from coder to reviewer/steward (C3) and from implementation to product decisions (C5). Bonferroni-adjusted α = 0.025.

Supported
Show item-level results
Itemn HM HIGH ± SDCI Hn LM LOW ± SDCI LTestStatisticpEffectVerdict
C387+1.02 ± 1.52[+0.70, +1.35]12-1.50 ± 1.09[-2.19, -0.81]Mann-Whitney UU = 934.5p < .001r = -0.790 (large)SUPPORTED
C587+1.24 ± 1.40[+0.94, +1.54]12-1.00 ± 1.54[-1.98, -0.02]Mann-Whitney UU = 889.5p < .001r = -0.704 (large)SUPPORTED

H1 and H4 produce large effects, but both compare 87 respondents against 12, and that imbalance is a direct consequence of the screening criteria. The effects are reported because they were tested, not because the design supports strong claims about their magnitude. H2 found no team-size effect at all — neither on the block mean nor on any individual item.

Open-text responses

What would need to change in your organization?

67 responses · 30 words on average

Themes
Team & Collaboration66
Organization & Culture47
Training & Skills25
AI Tools & Quality23
Governance & Guidelines12
Multiple assignment possible – one response can touch several themes.

Frequent terms

  • teams 27
  • team 19
  • system 17
  • organization 15
  • model 13
  • work 13
  • tools 12
  • shared 10
  • platform 10
  • training 10
  • people 10
  • change 9
  • company 9
  • working 9
  • structure 9

Frequent phrases

  • dedicated teams 4
  • platform teams 4
  • model work 3
  • individual tool 3
  • prompt engineering 3
  • product teams 3
  • team structure 3
  • shared knowledge 2

Selected responses

  • transitioning to this model requires shifting from individual tool adoption to a centralized platform engineering mindset where AI is managed as shared infrastructure rather than a personal assistant.
  • We would need to employ more people, our operations team is understaffed as it is, plus we don't have people experienced enough to design, implement and maintain such a system.
  • I think AI progress would have to slow down before we consider implementing this. With the current pace of innovation, it's better to let everyone have their own setup, at least for small startups like the one I work in.
  • I think that even two levels would be enough. In my organization, we are actively experimenting with a similar approach, and I really like it. The main challenge for organization is the budget (to have a separate team for creating and maintaining AI framework)
  • Integrating all members into a single system with shared AI tools would be the biggest challenge, since, as the company is small, it relies on each developer having the freedom to choose their tools, which would take some time to change the mindset
  • We would have to go slower to implement this - we don't have time to wait for dependencies between these teams

6 of 67 responses, selected to show the range of positions – supportive as well as sceptical. Reproduced unchanged, including typos. The full set of responses is not published.

Which agile practice has changed the most?

58 responses · 18 words on average

Themes
AI Tools & Quality40
Process & Agile28
Testing & CI/CD15
Team & Collaboration7
Training & Skills6
Multiple assignment possible – one response can touch several themes.

Frequent terms

  • code 22
  • changed 14
  • time 11
  • tools 10
  • development 9
  • work 9
  • faster 8
  • now 8
  • sprint 7
  • developers 6
  • agile 6
  • planning 6
  • since 6
  • reviews 5
  • rather 5

Frequent phrases

  • spend time 4
  • code review 4
  • code reviews 3
  • nothing changed 2
  • changed tools 2
  • generate code 2
  • reviewing validating 2
  • refining ai-generated 2

Selected responses

  • Code reviews have changed the most. I now spend less time on syntax and more time auditing AI logic for accuracy, security, and architectural fit.
  • I feel now the need for test driven development has increased a lot since AI has become a permanent part of a developer's life, since the code generated by AI needs a lot of evaluation and testing before it can be pushed to the production site.
  • The two-week sprint has become less meaningful due to the integration of AI.
  • Personally the team's agile practice remained the same even though all developers and PO use AI. It's more of AI helping us in bouncing off ideas and developers use to quickly create a quick POC for our clients. But agile methods remain the same for us

4 of 58 responses, selected to show the range of positions – supportive as well as sceptical. Reproduced unchanged, including typos. The full set of responses is not published.

Limitations

Imbalanced comparison groups. The screening criteria produced 87 intensive against 12 occasional AI users, and 88 small against 11 large teams. H1, H2 and H4 all rest on these splits. The effect sizes should be read as indicative, not as precise estimates.

Blocks are not scales. As set out above, the four blocks are thematic groupings. Block means summarise a theme; they do not measure a latent construct, and they should not be treated as scale scores in secondary analysis.

Self-report on a contested topic. All measures are perceptions. Whether respondents' roles have actually shifted from producing to reviewing code is not observed here — only that they report it.

Convenience sample. Prolific respondents who opt into a survey about AI tools are unlikely to be representative of developers generally, and the sample skews towards smaller teams and fewer years of experience.

Single measurement point. The data are cross-sectional and were collected in March 2026, in a fast-moving field. Several respondents noted in the open text that their answer would depend on how quickly the tools keep changing.

The instrument

The complete questionnaire as respondents saw it, exported from LimeSurvey. This includes the response options nobody selected — they were on offer, and that they stayed empty is itself a finding: no respondent identified as a Scrum Master or Agile Coach, none reported never using AI tools (the screening excluded them), and Tabnine went unused.

Two questions were asked but not analysed in this study: which agile frameworks the team uses (A5) and whether the team works with agile or lean methods at all (A6). Both served as screening and context.

AI-Dev-Systems: Organizational Design for AI-Augmented Software Development

Welcome!

Thank you for participating in this research study.

We are investigating how AI tools are changing software development practices and team structures. The survey takes approximately 7 minutes to complete.

Your responses are anonymous and will be used for academic research only.

Researchers: Andreas Hinderks (Hannover University of Applied Sciences), Joerg Thomaschewski & Eva-Maria Schoen (Emden/Leer University of Applied Sciences)

Screening & Demographics

Please answer the following questions about your professional background and experience with AI tools.

Prolific Participant ID

Free-text field

How many years of professional experience do you have in software development?

  • Less than 1 year
  • 1-2 years
  • 3-5 years
  • 6-10 years
  • 11-15 years
  • 16-20 years
  • More than 20 years

What is your current primary role?

  • Software Developer / Engineer
  • Software Architect
  • Team Lead / Engineering Manager
  • Scrum Master / Agile Coach
  • Product Owner / Product Manager
  • UX Designer / Researcher
  • QA / Test Engineer
  • DevOps / Platform Engineer
  • Full-Stack Developer
  • Other

How many people do you work with directly in your team?

  • 1-3
  • 4-6
  • 7-10
  • 11-20
  • More than 20

Which agile framework(s) does your team use? (Select all that apply)

  • Scrum
  • Kanban
  • Extreme Programming (XP)
  • SAFe
  • LeSS
  • Nexus
  • Shape Up
  • Other

Do you currently work in a team that uses agile or lean methods?

  • Yes
  • No

Which AI-powered development tools do you currently use? (Select all that apply)

  • GitHub Copilot
  • Cursor
  • Claude (Code / Chat)
  • ChatGPT
  • Gemini
  • Windsurf / Codeium
  • Amazon CodeWhisperer / Q
  • Tabnine
  • Other

How frequently do you use AI tools in your daily development work?

  • Never
  • Rarely
  • Occasionally
  • Frequently
  • Daily

How large is your organization (total employees)?

  • 1-10 (Startup)
  • 11-50 (Small)
  • 51-200 (Medium)
  • 201-1,000 (Large)
  • More than 1,000 (Enterprise)

Block 1: Perceived Changes from AI Tools

Please rate the following statements based on your personal experience with AI tools in software development.

Scale: 1 = Strongly Disagree ... 7 = Strongly Agree

  1. C1AI tools have significantly changed how I approach coding tasks.
  2. C2Since using AI tools, implementation is no longer the main bottleneck in my work.
  3. C3My role has shifted from writing code to reviewing and steering AI-generated outputs.
  4. C4AI tools allow my team to iterate much faster than before.
  5. C5I spend more time on product decisions (what to build) than on implementation (how to build it).
  6. C6Different team members use different AI tools, which creates friction.
  7. C7My team lacks clear guidelines on how to use AI tools effectively.
  8. C8Pair programming with AI feels fundamentally different from pair programming with a human colleague.

Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree

Block 2: AI Integration as an Organizational Challenge

The following statements address whether AI integration is primarily a tool adoption challenge or an organizational design problem.

Scale: 1 = Strongly Disagree ... 7 = Strongly Agree

  1. O1Integrating AI tools effectively requires changes to team structures, not just tool adoption.
  2. O2Framing AI as an "assistant" or "copilot" underestimates its organizational impact.
  3. O3Our organization would benefit from dedicated teams that design and maintain AI development environments.
  4. O4Each developer independently selecting their own AI tools leads to fragmented tooling across the team.
  5. O5AI tools should be managed as shared infrastructure rather than individual productivity tools.
  6. O6Governance (e.g., data privacy, quality standards) for AI-generated code is currently insufficient in my organization.
  7. O7Successful AI integration is primarily a technical challenge, not an organizational one. [reverse-coded]

Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree

Block 3: Three-Layer Model Plausibility

Please read the following description carefully, then rate the statements below.

Imagine that instead of each developer independently choosing and configuring their own AI tools, your organization creates a shared AI development environment — an "AI-Dev-System". This is a configured environment of AI agents, automated workflows, and curated knowledge bases that the entire organization uses for software development.

Three types of teams work together to make this system function:

System Teams design and build the AI-Dev-System. They select and configure AI agents, define interaction workflows, curate project-specific knowledge bases, and establish prompt engineering standards. Their goal is to encode organizational expertise into the system so that individual developers don't need deep AI configuration skills.

Platform Teams maintain and evolve the AI-Dev-System after deployment. They handle updates when foundation models change, fix issues that arise in daily use, extend capabilities based on team feedback, and manage regression testing. They ensure the system stays reliable and improves over time.

Product Teams develop software within the AI-Dev-System. rather than selecting AI tools individually, they consume curated capabilities provided by the platform — similar to how teams today use CI/CD pipelines without building them from scratch.

Scale: 1 = Strongly Disagree ... 7 = Strongly Agree

  1. M1The described three-layer model is a plausible way to organize AI-augmented software development.
  2. M2Having dedicated System Teams design AI environments would reduce the "paradox of choice" problem in my organization.
  3. M3Platform Teams maintaining AI-Dev-Systems would help my team work more consistently.
  4. M4As a Product Team member, I would prefer consuming a curated AI environment over selecting tools myself.
  5. M5This model would work in my current organizational context.
  6. M6This model requires a level of organizational maturity that most companies do not yet have.
  7. M7Small teams (fewer than 10 people) could implement a simplified version of this model.
  8. M8This model risks creating new silos between the three team types.

Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree

What would need to change in your organization to make a model like this work?

Free-text field · Please share your thoughts in a few sentences. Leave blank if you prefer not to answer.

Block 4: Agile Practices in Transition

Finally, please consider how AI tools are affecting established agile practices.

Scale: 1 = Strongly Disagree ... 7 = Strongly Agree

  1. A1The two-week sprint is becoming less meaningful as AI compresses implementation cycles.
  2. A2Test-driven development is becoming more important as a specification mechanism for AI-generated code.
  3. A3Retrospectives should evaluate the AI development environment, not just team processes.
  4. A4Continuous integration is more critical than before because AI-generated code requires stronger automated quality checks.
  5. A5The most important developer skill is shifting from coding to evaluating and directing AI outputs.

Response scale: 1 - Strongly Disagree · 2 - Disagree · 3 - Somewhat Disagree · 4 - Neutral · 5 - Somewhat Agree · 6 - Agree · 7 - Strongly Agree

Which agile practice has changed the most since you started using AI tools, and how?

Free-text field · Please share your thoughts in a few sentences. Leave blank if you prefer not to answer.

Response scale

All statements were answered on the same seven-point scale. It was recoded symmetrically for the analysis:

In the questionnaireIn the analysis
1Strongly disagree−3
2Disagree−2
3Somewhat disagree−1
4Neutral±0
5Somewhat agree+1
6Agree+2
7Strongly agree+3

Methodology notes

  • All Likert items use a 7-point scale, normalized to -3 to +3 (0 = Neutral)
  • Original scale mapping: 1 → -3, 2 → -2, 3 → -1, 4 → 0, 5 → +1, 6 → +2, 7 → +3
  • 95% confidence intervals computed using t-distribution (df = n-1)
  • Item O7 is reverse-coded – positive values indicate disagreement with the reverse statement
  • Distribution bars show percentage of responses per scale point
  • Interpretation thresholds: Strong Agreement (M ≥ +1.5), Moderate Agreement (M ≥ +0.5), Neutral (M ≥ -0.5), Disagreement (M < -0.5)
  • Cronbach's Alpha: internal consistency per block; α ≥ .70 acceptable, ≥ .80 good, ≥ .90 excellent
  • Item-total correlation: Pearson r between item and sum of remaining block items (corrected)
  • Exploratory Factor Analysis: PCA with Varimax rotation; Kaiser criterion for factor extraction
  • KMO measure of sampling adequacy: ≥ .60 required, ≥ .80 meritorious
  • Factor loadings ≥ |.40| considered salient for factor assignment

How to cite

In text

Hinderks, A., Thomaschewski, J. & Schön, E.-M. (2026). From Tool Adoption to Organisational Design: Research protocol for a survey on AI integration in agile software development [Research protocol, Version 1.0]. https://hinderks.org/studien/isd-2026-ai-integration

BibTeX

@misc{hinderks2026from,
  author       = {Andreas Hinderks and Jörg Thomaschewski and Eva-Maria Schön},
  title        = {{From Tool Adoption to Organisational Design: Research protocol for a survey on AI integration in agile software development}},
  year         = {2026},
  howpublished = {Research protocol, Version 1.0},
  url          = {https://hinderks.org/studien/isd-2026-ai-integration},
  note         = {Accessed: <date>}
}
Version
1.0
Licence
CC BY 4.0
Data
Aggregated results in full on this page. Raw data are not published; verbatim open-text responses only as a curated selection.

Version history

  1. Version 1.0First publication of the protocol.

Frequently asked questions

The four blocks are thematic item groups, not reflective scales. They were defined in advance along the propositions being tested rather than extracted afterwards from a factor analysis. A low alpha here therefore indicates multidimensionality, not unreliable measurement. Blocks O and M also contain deliberate counter-pole items that correlate negatively with the rest of their block. Interpretation consequently proceeds at the item and sub-cluster level, not via block sums.

Respondents consider the three-layer model plausible in the abstract (M1: M = +1.49) but doubt it would work in their own organisation (M5: M = +0.76). This gap of 0.73 scale points between conceptual endorsement and perceived deployability is the central finding of the study. It is consistent with M6, where respondents largely agree that the model presupposes a level of organisational maturity most companies do not yet have.

Only to a limited degree. The effect d = 1.63 comes from a comparison of 87 intensive against 12 occasional AI users. The small comparison group limits statistical power and makes the effect estimate unstable, even though the difference itself is pronounced. The same caveat applies to H2 and H4. The imbalance does reflect the population, however: in a sample screened for developers who use AI tools, occasional users are naturally rare.

No. This protocol contains the aggregated results: descriptive statistics per item, reliability figures, factor loadings and test results. Person-level raw data are not published. Of the open-text responses, only a curated selection appears, chosen to reflect the range of positions.