Human-AI Collaboration

How can human-AI collaboration exceed the sum of its parts?

Andy Weeger

Neu-Ulm University of Applied Sciences

November 13, 2026

Motivation

Generation is cheap.
Verification is not.

How can we structure our work with AI to ensure that the sum is greater than the parts, rather than less?

Exercise

Think about the last time you used an AI tool for university or work.

Did it save you time overall, once you count the time spent reviewing, correcting, and rewriting what it produced?

Discuss with a neighbor.

03:00

The light side

Execution velocity

AI compresses the time it takes to get from blank page to first draft.

  • Substantial completion-time reductions in routine writing and standard coding tasks (Merali, 2025; Noy & Zhang, 2023; Peng et al., 2023)
  • Rapid prototyping and cold-start knowledge retrieval
  • Reduced friction when summarizing unfamiliar material

The equalizer effect

AI narrows the gap between novices and experts, at least initially.

Disproportionate performance gains accrue to novices and bottom-quartile performers (Brynjolfsson et al., 2023; Cruces et al., 2025).

Individual divergent thinking

More ideas, further apart, faster.

Higher ideation fluency and greater semantic distance between ideas when individuals brainstorm with an LLM (Doshi & Hauser, 2024; Hubert et al., 2024).

Light side at a glance

Mechanism Changes Evidence
Execution velocity Faster drafts, faster prototypes, faster orientation Noy & Zhang (2023); Peng et al. (2023); Merali (2025)
The equalizer effect Novices and low performers gain the most Brynjolfsson et al. (2023); Cruces et al. (2025)
Individual divergent thinking More ideas, greater semantic distance Doshi & Hauser (2024); Hubert et al. (2024)
Table 1: Three well-replicated, immediate benefits of individual AI use

The dark side

Cognitive offloading and skill degradation

Faster answers today, weaker retrieval tomorrow.

The reviewer’s bottleneck and technical debt

AI writes faster than humans can review.

Fast generation produces high volumes of superficially correct but subtly flawed work (He, Agarwal, et al., 2026; He, Miller, et al., 2026); downstream, this shows up as silent coupling and structural complexity in code and analysis (Borg et al., 2026; Xu et al., 2025).

The fixation and homogenization trap

Everyone converges on the model’s first idea.

Reasoning biases and verification failures

Confidence is not correctness, but it is persuasive.

Dark side at a glance

Issue Mechanism Evidence
Cognitive atrophy Offloaded effort stops building the underlying skill Bastani et al. (2024); The Lancet Gastroenterology & Hepatology (2025)
Reviewer’s bottleneck Generation outpaces verification, debt accumulates silently He, Agarwal, et al. (2026); He, Miller, et al. (2026); Xu et al. (2025); Borg et al. (2026)
Fixation & homogenization Individual anchoring plus collective variance collapse Wadinambiarachchi et al. (2024); Doshi & Hauser (2024); Zhou et al. (2025)
Reasoning biases Automation bias and selective adherence to AI advice Vasconcelos et al. (2023); Alon-Barkat & Busuioc (2023)
Table 2: Four systemic penalties hidden behind the light side’s immediate gains

The jagged frontier

Mapping task performance

Figure 1: The jagged AI frontier, adapted from Dell’Acqua et al. (2023): AI ability (solid) does not track task difficulty (dashed) smoothly, so tasks of similar apparent difficulty can fall inside or outside the frontier

Exercise

Probe the frontier.

In pairs plus AI:

  1. Pick one real task from your own coursework or work that you’d consider delegating to AI. (2 min)
  2. Try it. Then actively try to break it, push an edge case, add an unusual constraint, or ask the model to show its work, looking for a spot where the output looks right, but you can’t actually tell if it is. (5 min)
  3. Report: what did you try, did you find a spot where the output was wrong or unverifiable, and what made that spot different from where it worked fine? (3 min)
10:00

What makes complementarity work

Verifiable and modular tasks are where humans and AI genuinely add up.

Determinant Q-uestion Significance
Verifiability Can an error be caught mechanically (linter, test, calculation check), rather than only by careful reading? Converts verification from an open-ended cognitive task into a fast, reliable one
Task modularity Can the sub-problem be fully isolated from the rest of the analysis? Bounds what the AI can silently get wrong, and what the human must review
Table 3: Two properties that predict when human-AI pairing outperforms either party alone (Hemmer et al., 2025; Vaccaro et al., 2024)

Interaction modes

From diagnosis to protocol

Four modes, matched to four failure risks.

  • Socratic Scaffold protects skill acquisition
  • Centaur Modularizer structures large workflows
  • Cognitive Forcer mitigates overreliance
  • Delayed Access preserves creative variance

The Socratic Scaffold

Cognitive protection for skill acquisition

The model is prohibited from giving direct solutions. It is prompted to act as an examiner: asking guiding questions, hinting at missing boundary conditions, checking the human’s mental model.

The Centaur Modularizer

Structural separation for large, multi-step workflows

The human owns strategy, problem architecture, and final synthesis. The AI is delegated bounded, non-overlapping subroutines, formatting, extraction, boilerplate.

The Cognitive Forcer

Adversarial review for high-stakes decisions

The human commits to an independent hypothesis in writing first. The AI is then prompted as an adversarial red team to find edge-case failures and unstated assumptions.

Exercise

Rewrite the moment

  1. Recall a specific time an AI tool gave you a confident answer that turned out to be wrong, or that you never actually checked. What was the task? What made you trust it? (3 min, individually)
  2. Share your story. Together, diagnose: could you have caught the flaw if you had formed your own view first? What was missing from how you used the AI in the moment? (5 min, in pairs)
  3. Draft the Cognitive Forcer rewrite: the one-paragraph hypothesis you would have written before consulting the AI, and the prompt that would have turned it into your adversarial reviewer instead of your source. (5 min, in pairs)

It would be great if a few share the most surprising or relatable story with the plenum.

13:00

Delayed AI access

Preserving variance in creative and strategic work

Enforce an initial period of unassisted divergent thinking. Introduce AI downstream, using diverse prompting personas to break out of central semantic distributions.

Challenges

You want to become a more deliberate collaborator with AI, not just a faster one? Here are three challenges that might help you along the way.

  • Level 1 (Reflect): Keep a 14-day AI reliance diary. Log every AI interaction relevant to your studies or work, noting where the task sat on the jagged frontier, whether your stance was passive or active, how long verification took, and any moment you noticed automation bias in yourself. Close with a written reflection on your personal dependency and verification bottlenecks.
  • Level 2 (Change): Over three weeks, commit to applying at least two of the four protocols above (Cognitive Forcer and Delayed Access work well together) across your coursework or projects. Compare your outputs against your usual habits: what changed in cognitive effort, in the maintainability of what you produced, and in what you actually remember afterward?
  • Level 3 (Grow): Design and run a small field experiment (4 to 8 people, a study group or a work team). Compare a human-only team, a naive-AI team, and a Centaur-style team on the same complex deliverable. Write up whether true complementarity appeared, whether technical debt accumulated, and whether the AI-assisted teams’ outputs converged more than the human-only team’s did.

Reading list

For digging deeper, I recommend the sources cited here, organized by theme.

Frontiers and productivity

  • The jagged technological frontier: Dell’Acqua et al. (2023)
  • Productivity effects of generative AI: Noy & Zhang (2023)
  • Generative AI at work: Brynjolfsson et al. (2023)
  • Generative AI and developer productivity (GitHub Copilot): Peng et al. (2023)
  • Scaling laws for AI-assisted knowledge work: Merali (2025)
  • Generative AI and education-based productivity gaps: Cruces et al. (2025)

Cognitive offloading and reliance

  • Generative AI and learning: Bastani et al. (2024)
  • Endoscopist deskilling after AI exposure: The Lancet Gastroenterology & Hepatology (2025)
  • Cognitive forcing functions: Buçinca et al. (2021)
  • Generation, verification, and AI-assisted decisions: Vasconcelos et al. (2023)
  • Generative AI and critical thinking: Lee et al. (2025)

Fixation and collective homogenization

  • Individual creativity vs. collective diversity: Doshi & Hauser (2024)
  • AI more creative than humans on divergent thinking: Hubert et al. (2024)
  • Generative AI and design fixation: Wadinambiarachchi et al. (2024)
  • Generative AI and divergent thinking in product design: Lin & Xie (2026)
  • The crowdless future: generative AI and collective problem-solving: Boussioux et al. (2024)
  • When ChatGPT is gone: creativity and homogeneity over time: Zhou et al. (2025)
  • Delayed AI access: Romero (2025)
  • Diverse AI personas and homogenization: Wan & Kalman (2025)
  • Diagnosing the limits of convergent AI: Williams (2026)

Reasoning biases and verification

  • Automation bias and selective adherence in public-sector decisions: Alon-Barkat & Busuioc (2023)
  • Automation bias in AI-assisted diagnostic reasoning: Qazi et al. (2025)
  • LLM influence on diagnostic reasoning: Goh et al. (2024)
  • LLM assistance and physician decision-making: Rounding et al. (2025)

Complementarity and downstream technical debt

  • AI-assisted teams vs. AI-led vs. human-only teams: Brodeur et al. (2026)
  • Field experiment on generative AI and teamwork: Dell’Acqua et al. (2025)
  • AI writes faster than humans can review: He, Agarwal, et al. (2026)
  • Speed at the cost of quality (Cursor AI): He, Miller, et al. (2026)
  • AI-assisted programming and developer productivity: Xu et al. (2025)
  • Downstream effects of AI assistants on maintainability: Borg et al. (2026)
  • The expert penalty in high-skilled work: Cui et al. (2025)
  • Early-2025 AI and experienced developer productivity: Becker et al. (2025)
  • When combinations of humans and AI are useful: Vaccaro et al. (2024)
  • Complementarity in human-AI collaboration: Hemmer et al. (2025)
  • Conversational AI for learning programming: MacNeil et al. (2024)

Q&A

Literature

Alon-Barkat, S., & Busuioc, M. (2023). Human-AI interactions in public sector decision-making: “Automation bias” and “selective adherence” to algorithmic advice. Journal of Public Administration Research and Theory, 33(1), 153–169. https://doi.org/10.1093/jopart/muac007
Bastani, O., Choi, C., Foster, D. P., Ghani, N., Kolter, J. Z., & Schwartz, Z. (2024). Generative AI can harm learning: Evidence from high school mathematics (w32630). National Bureau of Economic Research. https://doi.org/10.3386/w32630
Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the impact of early-2025 AI on experienced open-source developer productivity. https://doi.org/10.48550/arXiv.2507.09089
Borg, M., Hewett, D., Hagatulah, N., Couderc, N., Söderberg, E., Graham, D., Kini, U., & Farley, D. (2026). Echoes of AI: Investigating the downstream effects of AI assistants on software maintainability. Empirical Software Engineering, 31(2), 48. https://doi.org/10.1007/s10664-026-10889-1
Boussioux, L., Doshi, A. R., Hauser, O. P., & Hosanagar, K. (2024). The crowdless future? Generative AI and creative problem-solving. Organization Science. https://doi.org/10.63383/sxit9497
Brodeur, A., Valenta, D., Marcoci, A., Aparicio, J., Mikola, D., Barbarioli, B., Alexander, R., Deer, L., Stafford, T., Vilhuber, L., et al. (2026). AI-assisted teams outperform AI-led teams but not human-only teams in assessing research reproducibility in quantitative social science. Proceedings of the National Academy of Sciences, 123(8), e2524747123. https://doi.org/10.1073/pnas.2524747123
Brynjolfsson, E., Li, D., & Raymond, L. R. (2023). Generative AI at work (w31161). National Bureau of Economic Research. https://doi.org/10.3386/w31161
Buçinca, Z., Malireddy, M. B., & Gajos, K. Z. (2021). To trust or to think: Cognitive forcing functions can reduce overreliance on AI in AI-assisted decision-making. Proceedings of the ACM on Human-Computer Interaction, 5(CSCW1), 1–21. https://doi.org/10.1145/3449287
Cruces, G., Fernández Meijide, D., Galiani, S., Gálvez, R., & Lombardi, M. (2025). Does generative AI narrow education-based productivity gaps? Evidence from a randomized experiment (6246347). Social Science Research Network. https://doi.org/10.2139/ssrn.6246347
Cui, K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). The effects of generative AI on high-skilled work: Evidence from three field experiments with software developers. Management Science, 71(4), 2110–2135. https://doi.org/10.1287/mnsc.2025.00535
Dell’Acqua, F., Ayoubi, C., Lifshitz-Assaf, H., Sadun, R., Mollick, E., Mollick, L., Han, Y., Goldman, J., Nair, H., Taub, S., & Lakhani, K. R. (2025). The cybernetic teammate: A field experiment on generative AI and teamwork. Organization Science, 36(2), 485–507. https://doi.org/10.1287/orsc.2025.20702
Dell’Acqua, F., McFowland, E., Mollick, E. R., Lifshitz-Assaf, H., Kellogg, K., Rajendran, S., Krayer, L., Candelon, F., & Lakhani, K. R. (2023). Navigating the jagged technological frontier: Field experimental evidence of the effects of AI on knowledge worker productivity and quality (24-013). Harvard Business School. https://doi.org/10.2139/ssrn.4573321
Doshi, A. R., & Hauser, O. P. (2024). Generative AI enhances individual creativity but reduces the collective diversity of novel content. Science Advances, 10(28), eadn5290. https://doi.org/10.1126/sciadv.adn5290
Goh, E., Gallo, R. J., Hom, J., Strong, E., Weng, Y., Kerman, H., Cool, J. A., Kanjee, Z., Parsons, A. S., Ahuja, N., Horvitz, E., Yang, D., Milstein, A., Olson, A. P. J., Rodman, A., & Chen, J. H. (2024). Influence of a large language model on diagnostic reasoning: A randomized clinical vignette study. https://doi.org/10.1101/2024.03.12.24303785
He, H., Agarwal, S., Denisov-Blanch, Y., Azaletskiy, P. S., Koyejo, S., & Vasilescu, B. (2026). AI writes faster than humans can review: A longitudinal study of an enterprise 2x mandate. https://arxiv.org/abs/2602.09112
He, H., Miller, C., Agarwal, S., Kästner, C., & Vasilescu, B. (2026). Speed at the cost of quality: How cursor AI increases short-term velocity and long-term complexity in open-source projects. Proceedings of the 2026 ACM International Conference on Software Engineering (ICSE). https://doi.org/10.1145/3793302.3793349
Hemmer, P., Schemmer, M., Kühl, N., Vössing, M., & Satzger, G. (2025). Complementarity in human-AI collaboration: Concept, sources, and evidence. European Journal of Information Systems, 34(2), 180–199. https://doi.org/10.1080/0960085X.2025.2475962
Hubert, P. et al. (2024). The current state of artificial intelligence generative language models is more creative than humans on divergent thinking tasks. Scientific Reports, 14, 3440. https://doi.org/10.1038/s41598-024-53856-9
Lee, H.-P., Sarkar, A., Tankelevitch, L., Drosos, I., Rintel, S., Banks, R., & Wilson, N. (2025). The impact of generative AI on critical thinking: Self-reported reductions in cognitive effort and confidence effects from a survey of knowledge workers. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems, 1–17. https://doi.org/10.1145/3706598.3713778
Lin, X., & Xie, Y. (2026). Stimulating or constraining creativity? Traditional vs. Generative AI on divergent thinking in product design. Frontiers in Psychology, 17, 1184201. https://doi.org/10.3389/fpsyg.2026.1184201
MacNeil, S. et al. (2024). Exploring the design space of conversational AI for learning programming. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1–15. https://doi.org/10.1145/3613904.3642451
Merali, A. (2025). Scaling laws for economic productivity: Experimental evidence in LLM-assisted consulting, data analyst, and management tasks. https://arxiv.org/abs/2512.21316
Noy, S., & Zhang, W. (2023). Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654), 187–192. https://doi.org/10.1126/science.adh2586
Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). The impact of AI on developer productivity: Evidence from GitHub copilot. https://doi.org/10.48550/arXiv.2302.06590
Qazi, I., Ali, P. A., Asad, P., Khawaja, U., Akhtar, M. J., Sheikh, M. A., Muhammad, M., & Alizai, H. (2025). Automation bias in large language model assisted diagnostic reasoning among AI-trained physicians. https://doi.org/10.1101/2025.08.23.25334280
Romero, M. (2025). Delayed AI access: Mitigating the homogenising effects of generative AI in collaborative problem-solving. 2025 2nd International Conference on Artificial Intelligence and Teacher Education (ICAITE), 145–151.
Rounding, N. et al. (2025). Impact of LLM assistance on physician decision making: A multi-country randomized controlled trial. https://doi.org/10.1101/2025.08.08.25333272
The Lancet Gastroenterology & Hepatology. (2025). Endoscopist deskilling risk after exposure to artificial intelligence in colonoscopy. The Lancet Gastroenterology & Hepatology, 10(4), 312–314. https://doi.org/10.1016/S2468-1253(25)00133-5
Vaccaro, M., Almaatouq, A., & Malone, T. W. (2024). When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour, 8, 1892–1904. https://doi.org/10.1038/s41562-024-02024-1
Vasconcelos, H., Jörke, M., Grunde-McLaughlin, M., Gerstenberg, T., Bernstein, M. S., & Krishna, R. (2023). Generation meets verification: AI-assisted decision making with large language models. Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, 1–18. https://doi.org/10.1145/3544548.3581389
Wadinambiarachchi, G. et al. (2024). The effects of generative AI on design fixation and divergent thinking. Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems, 1–14. https://doi.org/10.1145/3613904.3642702
Wan, S., & Kalman, Y. M. (2025). Diverse AI personas can mitigate the homogenization effect in human-AI collaborative ideation. Computers in Human Behavior, 162, 108450. https://doi.org/10.1016/j.chb.2024.108450
Williams, R. (2026). In search of “weird corners”: Diagnosing the limits of convergent AI in professional creative practice. Proceedings of the 2026 Conference on Creativity and Cognition (c&c), 102–115.
Xu, F., Medappa, P. K., Tunç, M., Vroegindeweij, M., & Fransoo, J. C. (2025). AI-assisted programming decreases the productivity of experienced developers by increasing the technical debt and maintenance burden (5124892). Social Science Research Network.
Zhou, Y., Liu, Q., Huang, J., & Li, G. (2025). When ChatGPT is gone: Individual creativity fails to sustain while homogeneity keeps climbing. https://doi.org/10.5465/amproc.2025.16880abstract