GenAI for Critical Analysis: Practical Tools, Cognitive Offloading and Human Agency

I’m looking forward to LAK24 in Kyoto next month. In response to the GenAI workshop call for practical examples and experiences of GenAI, this short paper shares some examples from the last tumultuous year, with some brief reflections…

Buckingham Shum, S. (2024). Generative AI for Critical Analysis: Practical Tools, Cognitive Offloading, and Human Agency. In Joint Proceedings of LAK 2024 Workshops, co-located with 14th International Learning Analytics and Knowledge Conference, March 18-22, 2024, Kyoto, pp. 205-213. https://ceur-ws.org/Vol-3667/GenAILA-paper4.pdf

Abstract: Generative artificial intelligence (GenAI) is now capable of performing tasks that we have considered intellectually demanding. There are justified concerns that this will undermine the agency of both educators and students, if tools are poorly designed, poorly used, or imposed — with consequences for education and the future of work. This short paper contributes practical examples pointing the potential for GenAI to promote critical analysis as part of intellectually demanding tasks, by both students and educators. However, this depends on appropriate usage. The paper then briefly discusses how we may balance the benefits and risks of human cognitive offloading to AI, as a perspective on human agency.

Co-designing AI ethics in education

(Acknowledgements: DALL•E)

2024 here we come… The current frenzy around artificial intelligence in education was triggered just over a year ago by the explosive arrival of ChatGPT, which made the power of the most mature large language model ever developed, freely available to the masses via an engaging conversational user interface. Every level of the educational sector then spent 2023 grappling with the implications, and I’ve shared my small pieces of that puzzle in other blog posts. (For those interested, R&D in “AIED” is not new, dating back ~40 years depending on how you count.*)

The tech is advancing at a dizzying pace which can leave us disoriented, and few anticipate that 2024 will be any different. But a consistent challenge faced by every school, college and university, is to build and sustain trust that AI will be used responsibly. Easier said than done:

  • Local ethics. There are endless lists of AI ethics principles that would seem on first inspection to make sense everywhere (“fairness”, “accountability”, “transparency”, etc…) — but translation work is needed. There will be local sensitivities around how these are implemented. What do qualities like “trust” and “responsible” mean to teachers, students, parents, leaders, educational authorities? There will be commonalities for sure that translate across contexts, but building trust means taking your people on the journey, so that they can internalise what these ideas mean, bring abstract principles to life in their own language and metaphors, and tell user stories they can inhabit.
  • High quality deliberation. Moreover, the issues are complex. How do we convene informed, respectful dialogue between diverse stakeholders? Calling a ‘town hall’ for all interested risks being superficial (there’s no time to grapple with the complexities; contributions are misinformed), tokenistic (those in power have already made the decisions), or attracting only the most confident or strident voices. A brainstorming workshop provides more space to go deep, but often doesn’t involve any learning, participants may not represent the true diversity of the community, and while hugely generative of ideas, may fail to converge on tangible outcomes that actually make a difference.

Agreeing on what WE consider to be acceptable practice in OUR context can provide a sense of orientation and safety amid the turbulence — if they are then implemented of course.

In late 2021, here at UTS we set out to grapple with this, and designed the EdTech Ethics forum for the university community with these concerns in mind. We put out an initial report documenting the process and preliminary feedback in the immediate aftermath, but then did the key work of interviewing participants, analysing their feedback, followed by contributing to the university’s governance processes as it developed and published its AI Operations Policy and Procedures.

So I’m delighted to share a forthcoming journal paper documenting how we ran this, what the participants thought, and the tangible outcomes. The paper acknowledges the many people who made this possible, but special thanks to my co-authors Teresa Swist and Kal Gulson at Sydney University Education Futures Studio, who joined the project as external participants to UTS, and conducted the interviews. This work on Deliberative Democracy intersects with our collaboration around Technical DemocracyChad Foulkes from Liminal by Design was an awesome workshop session designer and facilitator, under the tricky lockdown conditions. And to Chris Riedy and Nivek Thompson (UTS Institute for Sustainable Futures) whose guidance and teaching on Leading Deliberative Democracy and Doing Deliberative Democracy started me down this road (highly recommended online micro credentials!).

Swist, T., Buckingham Shum, S. & Gulson, K. N. (2024). Co-producing AIED Ethics Under Lockdown: An Empirical Study of Deliberative Democracy in Action. International Journal of Artificial Intelligence in Education Published online: 27 Feb. 2024. https://doi.org/10.1007/s40593-023-00380-z

Abstract: It is widely documented that higher education institutional responses to the COVID-19 pandemic accelerated not only the adoption of educational technologies, but also associated socio-technical controversies. Critically, while these cloud-based platforms are capturing huge datasets, and generating new kinds of learning analytics, there are few strongly theorised, empirically validated processes for institutions to consult their communities about the ethics of this data-intensive, increasingly algorithmically-powered infrastructure. Conceptual and empirical contributions to this challenge are made in this paper, as we focus on the under-theorised and under-investigated phase required for ethics implementation, namely, joint agreement on ethical principles. We foreground the potential of ethical co-production through Deliberative Democracy (DD), which emerged in response to the crisis in confidence in how typical democratic systems engage citizens in decision making. This is tested empirically in the context of a university-wide DD consultation, conducted under pandemic lockdown conditions, co-producing a set of ethical principles to govern Analytics/AI-enabled Educational Technology (AAI-EdTech). Evaluation of this process takes the form of interviews conducted with students, educators, and leaders. Findings highlight that this methodology facilitated a unique and structured co-production process, enabling a range of higher education stakeholders to integrate their situated knowledge through dialogue. The DD process and product cultivated commitment and trust among the participants, informing a new university AI governance policy. The concluding discussion reflects on DD as an exemplar of ethical co-production, identifying new research avenues to advance this work. To our knowledge, this is the first application of DD for AI ethics, as is its use as an organisational sensemaking process in education.

Your thoughts welcomed in LinkedIn…

* Histories of AIED:

Doroudi, S. (2023). The Intertwined Histories of Artificial Intelligence and Education. International Journal of Artificial Intelligence in Education, 33(4), 885-928. https://doi.org/10.1007/s40593-022-00313-2

Pham, S. T. H., & Sampson, P. M. (2022). The development of artificial intelligence in education: A review in context. Journal of Computer Assisted Learning, 38(5), 14081421. https://doi.org/10.1111/jcal.12687

Woolf, B. P. (2015). AI and education: Celebrating 30 years of marriage. In C. Conati, N. Heffernan, A. Mitrovic, & M. F. Verdejo (Eds.), Artificial intelligence in education: 17th international conference, AIED 2015, Madrid, Spain, June 22–26, 2015. Proceedings (pp. 38–47). Springer International Publishing. https://doi.org/10.1007/978-3-319-19773-9

Assessment reform for the age of AI

“This is our Kodak moment”

This provocation was posed by one participant in a recent expert forum. Over-dramatic? Not if universities lose the capacity to assure learning.

Assessment Reform for the Age of Artificial Intelligence is a consultation report from the Tertiary Education Quality and Standards Agency (TEQSA), Australia’s independent national quality assurance and regulatory agency for higher education. I had the privilege of being invited to join a TEQSA-convened group for 2 days in August, which we hosted here at UTS, to consider how the assessment landscape has been redrawn by the emergence of widely available generative AI.

“The emergence of generative artificial intelligence (AI), while creating new possibilities for learning and teaching, has exacerbated existing assessment challenges within higher education. However, there is considerable expertise, based on evidence, theory and practice, about how to design assessment for a digital world, which includes artificial intelligence. AI is not new, after all. This document, constructed through expert collaboration, draws on this body of knowledge and outlines directions for the future of assessment. It seeks to provide guidance for the sector on ways assessment practices can take advantage of the opportunities, and manage the risks, of AI, specifically generative AI.”

As the report explains in setting the scene, the point of departure is a 2010 report entitled Assessment 2020: Seven Propositions for Assessment Reform in Higher Education, on which this new report is modelled [website/report]…

“We take our starting point for this document from the propositions for assessment outlined in Assessment 2020 (Boud and Associates, 2010). This work outlines how assessment acts as a powerful intervention in student learning and highlights the educational purposes of assessment in parallel with the process of assuring learning outcomes. Good assessment design that allows for ‘rich portrayals’ of student learning is critical. Thus, we take as given that assessment should engage students in learning, provide a partnership between teachers and students, and promote student participation in feedback. These key elements of assessment can then guide how best to consider the role of AI in assessment design.”

The question is — how does AI change things? We propose two propositions and five principles:

The lead team presented the report in this TEQSA webinar:

(This was the latest in a GenAI series co-hosted with Deakin University’s Centre for Research in Assessment and Digital Learning, to which I’ve had the privilege of contributing.)

Feedback is now in from the consultation, and this report will be presented and workshopped next month at the TEQSA conference, where we look forward to hearing more from participants. The question then, of course, is how to implement such changes at scale, in a sustainable way. As ever, Dave Boud has insights to offer…

It was fantastic working with such outstanding colleagues, and special thanks to the team who designed and facilitated the expert forum: Jason M. Lodge, The University of Queensland Sarah Howard, University of Wollongong Margaret Bearman, Phillip Dawson, Deakin University

With Shirley Agostinho, University of Wollongong, Simon Buckingham Shum, University of Technology Sydney, Chris Deneen, University of South Australia, Cath Ellis, The University of Sydney, Tim Fawns, Monash University, Helen Gniel, TEQSA, Rowena Harper, Edith Cowan University, Michael Henderson, Monash University, Danny Liu, The University of Sydney, Lina Markauskaite, The University of Sydney, Jan McLean, University of Technology Sydney, Carlo Perrotta, The University of Melbourne, Lambert Schuwirth, Flinders University, Christine Slade, The University of Queensland

Updates:

Understanding skilled use of open automated feedback tools as teacher feedback literacy

Summary: a new paper forges a bridge between data-driven, open automated feedback platforms, and teacher feedback literacy competences: 

Buckingham Shum, S., Lim, L.-A., Boud, D., Bearman, M. & Dawson, P. (2023). A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy. International Journal of Educational Technology in Higher Education, 20:40 (12 July 2023). https://doi.org/10.1186/s41239-023-00410-9 

The mass availability of generative AI continues to reshape thinking about the future of work and learning. Conversational apps can now give instant feedback to learners about their work — but the educational question is how effective this interaction is. A new design space has opened up for tuning generative AI to give high quality feedback to learners about their work. We are not in uncharted waters here: there is a growing body of knowledge on what “effective feedback” means in higher education, and how to create the conditions for this. It goes far beyond comments accompanying an assignment, with a shift towards “feedback rich ecosystems” in which both teachers and students exercise far greater agency and sensemaking competencies.

In 2019, an exciting book came out: The Impact of Feedback in Higher Education: Improving Assessment Outcomes for Learners (Eds. Henderson, Ajjawi, Boud & Molloy):

“This book asks how we might conceptualise, design for and evaluate the impact of feedback in higher education. Ultimately, the purpose of feedback is to improve what students can do: therefore, effective feedback must have impact. Students need to be actively engaged in seeking, sense-making and acting upon any information provided to them in order to develop and improve. Feedback can thus be understood as not just the giving of information, but as a complex process integral to teaching and learning in which both teachers and students have an important role to play. The editors challenge us to ask two fundamental questions: when does feedback make a difference, and how can we recognise that impact?”

In 2020, I conceived a symposium to bring the editors and authors to UTS to spend 2 days in dialogue with CIC and other researchers developing automated-feedback tools using Learning Analytics/AI. We called for a deeper dialogue between researchers in the design of assessment and feedback in higher education, and researchers developing automated-feedback tools using Learning Analytics/AI. The pandemic shifted this online, but the goals remained the same, and moving online enabled us to more easily bring in additional participants, resulting in DAFFI 2020: Designing Automated Feedback for Impact whose presentations I commend to you.

I’m now delighted to share one of the fruit from this, a collaboration between CIC (Lisa Lim and myself) and our colleagues at Deakin University’s Centre for Research in Assessment and Digital Learning (CRADLE). The focus of the paper is not on generative, conversational AI (which did not exist when we started this work), but on technically less complicated, but correspondingly far more transparent platforms that use simple rules authored by teachers themselves.

“In contrast to closed AF tools, we define open” AF tools as enabling the educator to specify some or all of the following key parameters in the tool’s behaviour:

  1. the student activity data that the system analyses;

  2. the algorithms that analyse that data;

  3. the feedback information the teacher wishes the software to compile for students;

  4. the modalities via which feedback information is communicated by teachers;

  5. the student-driven feedback processes that are afforded.”

What does it mean to do this skillfully? We demonstrate that Boud & Dawson’s  teacher feedback literacy competency framework can be applied very usefully to analysing teaching practices with data-driven, automated feedback platforms. A next step will be to think through what this means for tuning large language models for educational contexts.

A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy

Simon Buckingham Shuma, Lisa-Angelique Lima, David Bouda,b,c, Margaret Bearmanb, Phillip Dawsonb

a University of Technology Sydney, AUS
b Deakin University, AUS
c Middlesex University, UK

Effective learning depends on effective feedback, which in turn requires a set of skills, dispositions and practices on the part of both students and teachers which have been termed feedback literacy. A previously published teacher feedback literacy competency framework has identified what is needed by teachers to implement feedback well. While this framework refers in broad terms to the potential uses of educational technologies, it does not examine in detail the new possibilities of automated feedback (AF) tools, especially those that are open by offering varying degrees of transparency and control to teachers. Using analytics and artificial intelligence, open AF tools permit automated processing and feedback with a speed, precision and scale that exceeds that of humans. This raises important questions about how human and machine feedback can be combined optimally and what is now required of teachers to use such tools skillfully. The paper addresses two research questions: Which teacher feedback competencies are necessary for the skilled use of open AF tools? and What does the skilled use of open AF tools add to our conceptions of teacher feedback competencies? We conduct an analysis of published evidence concerning teachers’ use of open AF tools through the lens of teacher feedback literacy, which produces summary matrices revealing relative strengths and weaknesses in the literature, and the relevance of the feedback literacy framework.  We conclude firstly, that when used effectively, open AF tools exercise a range of teacher feedback competencies. The paper thus offers a detailed account of the nature of teachers’ feedback literacy practices within this context. Secondly, this analysis reveals gaps in the literature, signalling opportunities for future work. Thirdly, we propose several examples of automated feedback literacy, that is, distinctive teacher competencies linked to the skilled use of open AF tools.

Your comments most welcome

ChatGPT: What have we learnt, what do we need to learn next?

ChatGPT: What have we learnt?
What do we need to learn next?

I was honoured to join a TEQSA/CRADLE panel yesterday, the 3rd in a series on the implications of ChatGPT (or GenAI more broadly) for higher education. Nearly 3000 people registered, with >1200 joining live, reflecting either the gravity of the situation now facing us — or the consequences of AI and assessment becoming mainstream media fodder! It’s both in fact.

In the 2nd panel in March, in my 8min slot I flagged the absence (at that early stage) of any evidence about whether students have the capacity to engage critically with ChatGPT. So many people were proposing to do interesting, creative things with students — but we didn’t know how it would turn out.

But 3 months on, we now have:

  • myriad demos of GPT’s capabilities given the right prompts
  • a few systematic evaluations of that capability
  • myriad proposals for how this can enable engaging student learning
  • and a small but growing stream of educators’ stories from the field
  • with peer reviewed research about to hit the streets.

Educators can now articulate the range of critical engagement that their students are displaying, and I share what we’re learning at UTS from some of our leading educators who have been introducing assessments integrating ChatGPT. We now need to track how well these, and other interesting proposals, for AI-informed learning and assessment translate across diverse contexts.

I also urge us to harness the diverse brilliance of our student community in navigating this system shock, sharing what we’re learning from our Student Partnership in AI.

Here are my slides, and the full replay below (jumps to my 12min talk, but watch the whole panel!)

Conversational GenAI for argument analysis

History: making thinking visible

As you can tell from a quick scan of my site, I’ve spent a lot of time fascinated by how computers can make thinking visible (books Visualizing Argumentation and Knowledge Cartography), plus many papers and blog posts (on Argument Mapping and Dialogue Mapping), using experimental open source  software we built (such as Compendium and Cohere). A big picture account can be found in this 2007 keynote, Hypermedia Discourse: Contesting Networks of Ideas and Arguments.

So, all that’s to say that making arguments visible so that you and others can — in a very real sense — “see what you’re saying” has been a career-long passion. A key challenge in this long field of research has been that rigorous thinking is hard work. Bad luck, welcome to university! Argument Mapping and its related techniques use the affordances of visual trees/networks as an extended, external memory to augment personal and collective intelligence. Making one’s ideas visible as coherent diagrams is also hard work — but it’s good pain — the cognitive and discursive effort this entails is designed to clarify one’s thinking by revealing visually where the weaknesses are, in ways that writing and reading chunks of prose cannot tell you at a glance.

Enter NLP and rhetorical parsing

In 2012, we were now in the Web 2.0 era, and an exciting collaboration with NLP and linguistics expert Ágnes Sándor (Xerox) led to a new conception of Contested Collective Intelligence. For the first time in my work, machines could identify argumentative moves in sentences, complementing the argumentative moves that our web annotation tools enabled for people — who unlike machines, can of course can ‘read between the lines’ and see connections between ideas that may not even be in the texts.

This was extremely exciting, and the ideas and open source code carried through to our current Academic Writing Analytics project and web apps. I reflected  on the impact of encountering NLP colleagues, in the context of The Future of Text book.

Conversational generative AI

And so we arrive at generative AI based on large language models, which advances the state of the art in language processing and generation in so many ways. Moreover, the conversational paradigm, when a chat application is overlaid, opens so many interesting human-computer/personal-collective intelligence possibilities. I’ve been intrigued to play with GPT-4 to see what its argument analysis capabilities are.

Previously, I’ve shared some early experiments on ChatGPT-3.5’s ability to identify implicit premises in prose arguments, and critique a flawed argument by analogy. I’ve now had the chance to experiment a little with the version of GPT-4 that is Bing Chat, accessed via Microsoft Edge browser. I was dying to see how far I could get in generating an Argument Map from a written argument.

The task is a typical analysis workflow, as prep for teaching:

  • search for relevant sources
  • select one for analysis
  • extract key elements of the argument and their relationships (described using a structured markdown notation called ArgDown)
  • diagram them to show their key relationships (in the ArgDown web app)
  • discuss (with the AI)
  • start thinking about student activities to help them learn

I don’t mind admitting that watching a machine do this for the first time was startling! I tell the story here…

U21 2023 Educational Innovation Symposium Keynote from McMaster University (OFFICIAL) on Vimeo.

Deeper dive

Let’s take a closer look at what Bing Chat did, because it wasn’t perfect.

  • The gold stars signal what in my view are good summaries of what the authors said, correctly linked.
  • The blue info circles are “commentary” from Bing Chat about the arguments
  • The red crosses signal that the authors did not say this, it is a false reconstruction by Bing Chat.
  • The red underline signals classification of a premise using incorrect, or indeed made-up argument schemes. There is to my knowledge no such argument type as Argument from responsibility, or Argument from precaution. Argument from omission seems to be a jumbling of Fallacy of omission and Argument from ignorance. 

If we take this node for example, it reads well as a summary:

However, the authors do not talk about researchers at all, they say:

“The letter addresses none of the ongoing harms from these systems, including 1) worker exploitation and massive data theft to create products that profit a handful of entities, 2) the explosion of synthetic media in the world, which both reproduces systems of oppression and endangers our information ecosystem, and 3) the concentration of power in the hands of a few people which exacerbates social inequities.”

As an amusing sidenote, Bing Chat was curiously resistant to recognising this, insisting that it was correct, first “quoting” a fabricated passage from the article to me, and then saying that this implied that the authors meant researchers. I thought that this sort of stubbornness had been ironed out after Bing Chat’s earlier escapades! More seriously, this points to the value of dialogic learning, with a partner who can be conversed with 24/7 — but who must still be treated with some caution, certainly at this stage of maturity.

To summarise:

  • Bing Chat showed intriguing capability, for a machine, to analyse an argumentative article:
    • extracting the key claim and underlying premises, summarising them in own words
    • generating markdown (ArgDown) showing supporting/challenging relationships
    • (and without being asked to) attempting to classify some nodes using Walton’s Argumentation Schemes.
  • However it also introduced fallacious nodes (incorrect summaries of the authors, and incorrect commentary nodes), incorrect links, and argument classifications (inventing argument types, and/or misclassifying nodes).

This is an exploratory example, and more systematic evaluations are required, of the sort we see in the growing Argument Mining literature.

Reflections

It does feel to me that we’ve turned a corner in the long, wintry history of AI. Perhaps this is a passing summer, which will fade like the others. But in my own career, punctuated by eureka moments such as seeing my first Apple Mac, my first web page load, and an iPhone — this is up there.

University is to teach you to think. Argument analysis is serious intellectual work, of the sort that we would hope to see from our students. Nor is there always “one map to rule them all’ — a correct map, since like in spatial cartography, design decisions are made about scale and purpose. The point about knowledge cartography is that it provokes productive reflection and discourse. So even if the AI gets the map wrong (and it will), the conversation this should provoke should be useful. With colleagues Kirsty Kitty and Andrew Gibson, I’ve argued that embracing imperfection in tech can be productive if it promotes deeper critical thinking in learners, e.g., learning by correcting the automated output, or reflecting on questions it asks, or why it seems wrong. Students must, however, be scaffolded to engage in such activity.

Informal learning? This is feasible in formal education, but may be less attractive in other informal learning contexts where we want to promote critical deliberation, e.g.  citizens engaged in a policy deliberation, many of whom lack the internal or external motivation to think that hard. But assuming future tools give more accurate argument maps/outlines, that require less debugging, perhaps we can see use-cases including:

  • assisting facilitators/educators to prepare learning resources for civic deliberations
  • assisting very engaged citizens to dissect complex arguments, and perhaps lowering the entry threshold for others who might otherwise not engage with such structured, critical deliberation
  • an article is very different to a multi-author conversation, but we can envisage summarising online discussions (NB: Teams is starting to summarise topics and actions in meeting transcripts)

Did we just supplant student cognition? From a learning sciences perspective, an overriding concern with generative AI is that it does too much cognitive work for the learner. Editing an AI-generated draft is not the same as wrestling with the blank page yourself. Ditto for reviewing an AI-generated argument map.

I have just done what many professionals have enjoyed doing in recent months: putting GPT through its paces to test its technical capability. But learners are not professionals: they don’t know what they don’t know. As I argue elsewhere, they may lack the knowledge, skills and dispositions to engage critically with AI output. They will require suitable scaffolding from mentors and teachers to learn what we mean by critical thinking and argument analysis, in order then to be equipped to use a power tool such as an argument mapping tool. Much empirical research awaits to test the affordances of generative AI like this, to establish when they are most useful to use developmentally, with a given age/stage of learner.

But we do know that argument mapping has struggled to gain traction (in formal education and among professionals) because it’s hard intellectual work. It could be that by generating full or intentionally incomplete argument maps, AI provides a step up for many learners to quickly get feedback on their work, or see examples of arguments about topics they are knowledgeable about — and thus better equipped to critique — compared to examples chosen by the teacher or textbook. Generative AI may open new possibilities because it can generate examples tuned to the interests of each learner, activating their curiosity to go deeper.

Your feedback is welcome, which is hosted on LinkedIn…

The Writing Synth Hypothesis

The Writing Synth Hypothesis

Reflecting on where writing is heading seems critical as, within education, we think about the future we should equip our graduates for, which in turn should shape the future of writing pedagogy and assessment.

An AI-generated image from DALLE•E showing sliders and knobs in a futuristic writing app

The hypothesis

Synthesisers transformed music composition fundamentally. As personal computers became widespread, and digital audio workstations with built-in instruments and effects became affordable in the late 1980s, the masses could start tinkering with audio tracks without needing to learn an instrument or the formal fundamentals of music. Non-linear editing was a fundamentally different way of composing, enabling the flexible exploration of creative options.

The Writing Synth hypothesis proposes that with the emergence of generative AI, authors will be able to learn writing in new ways, democratising writing just as we saw with music synthesisers.

Now we need to learn to play these new instruments.

There may be new genres of writing that, like the music revolution, were impossible to create without these new tools.

This is a working hypothesis.

  • It needs to be tested conceptually (does the argument by analogy hold up?).
  • There’s important user interface design work to do (since writing is different to music, how will we orchestrate texts?).
  • And the vacuum of evidence must be filled (what does such writing look like in practice, who is capable of it, and does it assist learners of all ages and stages?).

Let’s take a walk to explore this new space.

Music synths: data, interoperability, UX

The music synth revolution was possible thanks to a radical new data and interoperability infrastructure. Analogue and then digital synthesisers could be connected to computers thanks to the new underlying standard for digitising and transmitting audio signals between devices called MIDI (Musical Instrument Digital Interface). But a data infrastructure is only useful when humans can interact with it, which brings us to the user experience (UX).

Younger readers will not recall a time before visual text editors. The ability to translate thoughts onto the screen fast enough to keep pace with one’s thinking was a revolution, first demonstrated in 1968 by Doug Engelbart in his extraordinary Mother of all Demos. We can barely conceive how revolutionary it was in the era of the typewriter and tickertape, to see someone type something, change their mind, and instantly edit it. With the PC revolution led by Xerox, Apple and Microsoft, we moved from command line interfaces (where the user had to type arcane command syntax and semantics) to what were first termed WIMP (Windows/Icons/Menus/Pointer) Graphical User Interfaces (GUIs), which we of course now take for granted. These displays were revolutionary, constantly reminding the user what commands were available via icons and menus (exploiting human recognition instead of recall), offered complementary, interlinked views of data (in these wonderful new windows) which could be arranged on screen, with myriad interactive ‘widgets’ such as checkboxes and sliders to set preferences.

The arrival of non-linear editors orchestrated these fundamentally new ways of interacting with digital assets to transition musical composition into playful experimentation with interactive, visual, multitrack timelines. Now, like text, audio edits had “undo”, and clips could be dragged+dropped, copied+pasted, merged+split, and ‘formatted’ by tweaking their many audio properties. Video followed closely behind, once computing hardware caught up to handle storage and resolution challenges.

Envisioning the AI Writing Studio

As someone coming from the Human-Computer Interaction (HCI) community, I am drawn to design prototypes as one way of envisioning the future, so we’ll kick off with that. The chat interface in OpenAI’s ChatGPT has seized the world’s imagination with its simplicity, providing the first walk-up-and-use interface to the large language model capability (which had been available for several years via GPT APIs, but only to technical experts). Everyone knew what textchat was, and it reinforced the conversational metaphor that played to the public’s sci-fi imaginations. The addition of voice input and output consolidated that narrative — at last AI had delivered HAL, C-3PO, DATA and all our other favourites from the movies. Our baby AI can talk (apparently about anything, with great confidence) and we’re absolutely besotted!

But when we remind ouselves that the user interface is a way to control a powerful computer, a chat metaphor is not the only, or even optimal, way to perform all tasks. In one sense, it’s a variant on the good old command line interface that preceded GUIs, requiring the user to know, like a magician conjuring spells, the commands that will invoke the most powerful effects. Those who have reached moderate to expert levels of proficiency with the Unix command line revel in the power this brings to control in ways that are impossible one click at a time in a GUI. We see the rise of this new art as people delight in figuring out ways to make ChatGPT do their bidding, and set themselves up as Prompt Engineering gurus. This is fun while we all play — but if you need to do serious work, the idea that you need to approach your AI assistant with guile and cunning — as though they’re a tetchy colleague you have to manipulate to get them to cooperate — seems odd to say the least.

While learning to control the output of language models is certainly a form of AI literacy, the need for “prompt engineering” may be consigned in the history books to a curiosity associated with the earliest releases, as people sought to use the chatbot not just for conversation, but as a practical creative tool. A command line interface with highly unpredictable output is not the optimal user interface for co-creation.

How might the UX evolve ? Firstly, taking inspiration from the music revolution, I anticipate the emergence of writing environments will enable authors to orchestrate their writing in new ways. Perhaps the introductory user guide to an AI Writing Studio (Sept. 2023 Release) will describe functionality like this…

  • Source Apps. Select which AI writing generators you want to work with — the studio will render their drafts in different windows which you can arrange, refreshing them each time settings are changed. After a while you may figure out which ones work best with different styles, which ones are most responsive, or which ones are most fun!
  • Genre menu. Choose the genre of writing you want to work in (e.g., tech blog; journal article; business report; etc.)
  • Modulators. Configure the libraries of sliders down the side that you want — these remind you how the text can be modified and encourage experimentation, applied either to the whole document, or the selected text (e.g., length; formality; reader age; etc.)
  • Record On/Off. Turn on recording to log your studio session, enabling Replay and Analytics (see below).
  • Replay. Fast Forward/Rewind through your document’s timeline to revisit key moments (e.g., recovering the state of the modulators and each app’s draft at a given moment — you can branch your document and explore another version).
  • Analytics. AI is used to generate summaries of your usage of AI generators, such as how much you request, reject, adopt or adapt AI suggestions. This helps evidence your critical engagement with AI, which your course will have mentioned. Check if your assignment requires you to include the WAL (Writing Analytics Link) and/or the 1-page report.
  • Feedback Tips. The feedback panel uses the best research on writing to assist your writing skills. In addition to the Analytics, other tabs show you well established indicators of the clarity of writing, and the depth of your reflection and argumentation (varies with the genre you chose). The analytics are your springboard into our personally recommended Practice Exercises and Pro Tips…
  • Practice Exercises. These help you get the most out AI writers, while ensuring that you’re building your own writing and thinking skills.
  • Pro Tips. We curated some of the best videos from our elite writers, who walk through their writing practices with Writing Visual Studio.

At some point perhaps I’ll mock the interface up, and even get to build it. But these are just preliminary ideas — there are far more creative possibilities, introduced next.

Moving beyond AI as ghostwriter demands creative UX design

Glenn Kleiman helpfully discusses (with the aid of GPT) the roles that AI writers can play — as editor, co-author, ghostwriter, and muse. The panic around cheating focuses on AI as ghostwriter, and we will watch the inevitable arms race between AI generators and detectors play out. Policing is important, but not the only mindset we need to adopt. What might interaction with AI writers playing the other roles look like?

We find clues in the communities spanning both academia and the tech industry who’ve been working on Computational Creativity and most recently, Human-AI Co-Creation with Generative Models. Consider Ken Arnold and colleagues, prototyping generative AI that augments rather than automates human writers. The purpose of the AI is to prompt the author with questions, promoting more critical thinking and better writing:

A team at MIT and Harvard are exploring the addition of audio and visual cues for creative writers, along with text: this paper exemplifies the kind of detailed analysis of tool usage that the Writing Synth hypothesis requires:

Continuing with creativity, consider the extraordinary work of Sarah Schwettmann who shows in this keynote talk how generative AI for The Met renders images of cultural artifacts that fill the gaps between artifacts that have been discovered. Thus, using “Generist Maps“, we can ask what might have happened if two cultures had met?

If this is possible for images, then is it not plausible to envision AI writers drafting texts in the interstitial spaces between known evidence, or between polarised positions in an argument? And might visual interfaces such as the above not be an engaging way to work with drafts ?

In another project, Schwettmann’s team describe the Latent Compass, a prototype to help individuals generate images that match their intuitions about the meaning of complex terms (e.g., “more festive”, “more inviting”).  In a writing context, I can see authors teaching their AI assistant with examples of what they mean by “the crisis motivating a research program”, or “artful critique of an argument”, which become new, user-defined modulators. I would see such developments as evidence for the writing synths hypothesis.

The Human-AI co-creation community also offers us conceptual language to describe how the human and AI can be configured to co-create together, such as this example from Michael Muller and colleagues:

[Update 27 Mar 2023] Most recently they have proposed a set of user-centred design principles to guide generative AI tools, which I am looking forward to thinking through in relation to the writing studio concept:

So, those are glimpses of where we may be heading with human-AI co-writing. But right now, in the true Silicon Valley ethos of “move fast and break stuff”, we are witnessing the largest scale introduction of AI in education, with no evidence of its utility for learning. And it is in the trenches of everyday education, at school and university, and professional learning in the workplace, where the writing synth hypothesis must be tested.

Teaching and assessing writing: many hopes and fears, little evidence

GenAI is a system shock because teaching and assessment regimes rest on the assumption is that the learner has written the text, and that the goal is to assess their ability to do so unaided by anyone or anything else (other than the passive capabilities of word processors). Learners may, of course, draw on others’ work, but only following well-established guidelines (e.g., through quoting and citing), in order to maintain academic integrity. Some students cross the line into the territory of student misconduct (the reasons for which are complex), and an array of policies and software products to police this are in place.

However, rather than simply banning AI writing, this has also triggered an outbreak of creativity as educators share and debate ways to actively embrace the new possibilities of composing with AI, as a new way to cultivate students’ critical faculties. This is in my view absolutely the way forward. Our graduates must know how to orchestrate these instruments and (to borrow an aviation metaphor) fly them within their ‘flight envelope’ — understanding the limits within which they can be trusted to perform reliably, before the wings drop off…  Beyond that, if they’re to find work in the creative professions, students must be able to show the additional value that distinguishes them from 100% AI-generated writing, or mere AI app operators who can simply click buttons.

In my own university we are advising effective ethical engagement, and resourcing academics for shorter and longer term adaptation of their assignments, and most other universities are doing the same. This is a holding pattern while we wait for the dust to settle. The web is full of proposals for engaging students in using ChatGPT creatively (101 ideas), while others warn of the death of thinking. These myriad hopes, fears and advice are filling the vacuum of evidence at this transition point. In a year’s time we’ll have many anecdotes and practitioner reports, and the first robust peer reviewed research evidence.

However, while generative AI is undeniably new, we are not in completely uncharted waters. AIED research has been under way for over 40 years. There are communities dedicated to prototyping and evaluating computational support for writing, conversational user interfaces and pedagogical agents, to name just three at the intersection of ChatGPT as a design concept. The media conversation would be more informed if these researchers can translate their work into accessible forms for wider audiences, as well as apply their expertise to show how generative AI can be designed and deployed in ways that respect with what we already know. We’ve made a start on that conversation in my own institution.

[Update 14 Aug 2023: An expert forum on the future of assessment in the age of AI just wrapped up, and the report will be shared in this TEQSA webinar]

Knowledge, skills and dispositions for critical engagement with AI

Amidst all the excitement among the optimists, let’s consider one of the most prevalent aspirations: that students will critically engage with AI draft writing, identify its weaknesses, and show how they have improved on it in their submitted work. While academics proposing these ideas are able to do this, I wonder if they overestimate their students’ knowledge, skills and dispositions to do so.

  • Curriculum/domain knowledge is needed to validate factual claims and spot significant omissions.
  • Rhetorical analysis and writing skills are needed to improve on prose which may in fact exceed many students own ability
  • Dispositions such as the curiosity and authenticity are needed to resist the temptation to just run with what the AI served up.

These qualities must be demonstrated rather than assumed, and educators should design for wide variability among their students in their capacity to critically engage. This is just one example of the evidence that needs to be gathered. [Update 8 June 2023: after 1 semester teaching with ChatGPT, we have initial evidence]

[Update 27 Mar 2023] The Academic Integrity debate around generative AI is (understandably) skewed to the ‘dark side’, and badly in need of more sophisticated vocabulary to talk about what it means to write with integrity with AI. Katy Gero’s exciting research illuminates how creative writers feel about AI writing aids, and is exactly the kind of work we need now. I’ll highlight just one aspect of her work, around the differing ways that writers feel about “authenticity”:

“Writers talked about authenticity, or their ‘voice’, as a concern when it came to incorporating the ideas or suggestions of others. Here, we describe four types of authenticity issues that came up in our interviews: 1) the reader’s sense of authenticity, 2) the impact of viewing suggestions, 3) differing opinions on where authenticity lies, and 4) human v. computer authenticity issues.” (Gero, 2022: p.106)

Gero’s work may offer us concepts and language to help students develop their own sense of what it feels like to work authentically with AI writers.

Writing analytics and academic integrity

(An earlier version of this section was originally posted here)

Recall Analytics in my envisioned writing studio. In the near future, GenAI will be fully integrated into interactive tools for writing, coding, and other creative work with image, music, animation, video etc… I envisage our students will become power-users. Human-AI interaction ‘flow states’ will become a synergistic blur, as prompts are invoked by the learner or offered by the machine, and rejected, adopted, adapted — each in the space of a few seconds. Tens of thousands of times in the production of an assignment.

Asking a student to “declare/document what role AI played” after hours/days/weeks of working in close partnership with such tools now becomes an impossible question to answer.

Instead, following the Writing Synth hypothesis, we look to the music world and borrow a studio recording session analogy: we immerse ourselves in our work, it’s all being recorded, and then we need to replay and review, dissect and debate, re-record elements, or start over…

In educational terms, such tools will be scaffolding “reflection-on-action” (Donald Schön) by the learner, possibly also with peers, and the teaching/coaching team. In time, they develop the capacity to engage in increasingly nuanced “reflection-in-action”, making improvised decisions about how and when to call on AI…  In the language of human-computer interaction research, such tools will support Retrospective Cued Recall.

Analytics crunching that data will make visible patterns that are useful for improving performance. My colleague Antonette Shibani has already prototyped this (see below). We will be able to see — literally — how virtuoso performance with such tools differs from less developed performances. This can serve as formative feedback to the learner, and assist should academic integrity questions arise.

[Update 26 Apr 2023] Antonette Shibani, et al (2023). Visual representation of co-authorship with GPT-3: Studying human-machine interaction for effective writing. 16th International Conference on Educational Data Mining

The fundamental question, then, is whether students are learning to produce great work. And in the future, great work will not be merely what can be automated. As Michael Feldstein has noted, students must learn the limits of GenAI, so that they develop the qualities needed to produce work that is beyond full automation — and stay employed.

And so we return to assessment.

If you can’t write without AI, can you really write?

In a prescient paper written at the turn of the 90s, Gavriel Salomon, David Perkins and Tamar Globerson considered critical educational questions that they envisaged arising with “intelligent technologies” as they termed them. When we ask what effect AI has on students, they distinguish between performance with the AI, and the effects of using AI on the student, assessable once the AI is removed. Intriguingly, they invite us to imagine a positive, futuristic scenario:

“For another illustration, consider the possible impact of a truly intelligent word processor: On the one hand, students might write better while writing with it; on the other hand, writing with such an intelligent word processor might teach students principles about the craft of writing that they could apply widely when writing with only a simple word processor; this suggests effects of it.”

Well, here we are! Fast-forward 30 years to today, and some have argued that ChatGPT is an educational disaster because we only learn to think by writing (Rob Reich, p.20). Decades of research into writing does indeed show that the writing process activates many cognitive faculties for critical thinking. But the roles that a conversational, generative AI agent can play in provoking deeper thinking (see above examples) are not taken into account by such cognitive models, which assume a solo author.

Salomon et al. argue for mindful versus mindless engagement with AI to achieve high performance with AI, and pose the assessment question now confronting us today: should we evaluate what a student is capable of when using AI to augment their intellect, or the “cognitive residue” as they term it — how well they perform once stripped of the AI ? For many educators, it would be a dereliction of duty to turn out graduates who could not write well with a pen and paper, while for others, that is to be stuck in the past. The imperative is to graduate capable of high performance with a profession’s state of the art tools. It may of course be a false dichotomy if the latter is impossible without the former, but that is an empirical question.

Writing in the early 90s, pre-Web, pre-mobile, pre-Big Data, and pre-LLMs, the authors conclude that we cannot afford to assess only AI-augmented student performance. After all:

“Until intelligent technologies become as ubiquitous as pencil and paper—and we are not there yet by a long shot—how a person functions away from intelligent technologies must be considered. Moreover, even if computer technology became as ubiquitous as the pencil, students would still face an infinite number of problems to solve, new kinds of knowledge to mentally construct, and decisions to make, for which no intelligent technology would be available or accessible.”

We might question this assumption now — but something deep inside us as educators might whisper that we will have really lost the plot if our graduates cannot function without computational support. The resolution may lie in what exactly we want students to bring. Rose Luckin and Margaret Bearman have argued that it is pointless to assess students on anything that AI can do better, which is a rapidly rising waterline (Salomon et al. contest this). I’ve also argued that we need to move to higher ground and  harness analytics and AI to help where they can in cultivating the qualities and capabilities that are still distinctively human. How about we start with dignity, compassion and justice.

To close…

So, that’s the Writing Synth Hypothesis. I had fun writing it — let’s see how it all unfolds. This is indeed an extraordinary time.

Your comments are most welcome: my blog doesn’t have great discussion tools, so join the conversation in this LinkedIn thread.

Framing Generative AI as EdTech

So far much of my year has been dominated by the widespread availability of generative AI apps, especially ChatGPT given my work in writing analytics. It’s been hectic but interesting connecting across the university, working closely with Kylie Readman (VP Education & Students) and my IML colleagues, to help prepare briefings and policy.

If you’re helping your institution develop responses to this, or are wondering as a researcher in EdTech/Learning Analytics/AIED how to engage, then you may be interested in:

Framing Generative AI as EdTech

1 hr UTS webinar, 23 February 2023

Simon Buckingham Shum is a Professor of Learning Informatics & Director, Connected Intelligence Centre.

Baki Kocaballi is a Senior Lecturer in the School of Computer Science. He is actively researching Conversational Interfaces and Human-AI Interaction.

Shibani Antonette is a Lecturer in the TD School, and actively researching Automated Writing Feedback and AI tools for education.

Generative AI (GenAI) is being hailed as a tipping point in AI, but let’s be clear: when it comes to educational technologies (EdTech), we have not just landed on “terra nullius”. While apps such as ChatGPT and DALL-E were never developed explicitly as EdTech, like so many other interactive tools we use every day, that doesn’t mean they have no educational value when used well. It’s too early to have peer-reviewed evidence of ChatGPT’s educational effectiveness, but prior research in related areas offers both theory, evidence and practice. So in this session, we’ll locate ChatGPT in the broader research landscape. Experts in two key fields will share their work on Automated Writing Evaluation, and Conversational Interfaces, where ChatGPT sits right at the intersection. Sharing brief glimpses of this work, we aim to spark ideas around how we can build on such foundations, to promote effective, ethical engagement with GenAI and avoid going down dead-ends that are already known from pre-GenAI research. [slides]

Human/AI exam proctoring with integrity?

Enjoyed convening this SoLAR Panel with some very knowledgeable colleagues…

The emergence of online exam proctoring (aka remote invigilation) in higher education may be seen as a function of multiple interacting drivers, including:

  • the rise of online learning
  • emergency exam measures required by the pandemic
  • cloud computing and the increasing availability of data for training machine learning classifiers
  • university assessment regimes
  • rising concerns around student cheating
  • accountability pressures from accrediting bodies

Commercial proctoring services claiming to automate the detection of potential cheating are among the most complicated forms of AI deployed at scale in higher education, requiring various combinations of image, video and keystroke analysis, depending on the services. Moreover, due to the pandemic, they were introduced in great haste in many institutions in order to permit students to graduate, with far less time for informed deliberation than would have been expected. Consequently, there was significant controversy around this form of automation, with protests at some universities seeing withdrawal of the services, and research beginning to clarify the ethical issues, and produce new empirical evidence.

However, numerous institutions are satisfied that the services they procured met the emergency need, and are continuing with them, which would make this one of the ‘new normal’ legacies of the pandemic. Critics ask, however, whether this should become ‘business as usual’. Regardless of one’s views, the rapid introduction of such complex automation merits ongoing critical reflection.

SoLAR was delighted to host this panel, which brought together expertise from multiple quarters to explore a range of questions, arguments, and what the evidence is telling us, such as…

  • This is just exams and invigilation in new clothes, right? They’re not perfect, but universities aren’t about to drop them anytime soon, so let’s all get on with it…
  • Are there quite distinct approaches to the delivery of such services that we can now articulate, to help people understand the choices they need to make?
  • What ethical issues do we now recognise that were perhaps poorly understood 2 years ago — or simply couldn’t afford to engage with in the emergency, but which we must address now?
  • What evidence is there about the effectiveness of remote proctoring — automated, or human-powered — at reducing rates of cheating?
  • What answers are there to the question, “Should we trust the AI?” Are we now over (yet another) AI hype curve, and ready for a reality check on what “human-AI teaming” looks like for online proctoring to function sustainably and ethically?
  • What (new?) alternatives to exams are there for universities to deliver trustworthy verification of student ability, and what are the tradeoffs?
  • Who might be better or worse off as a result of the introduction of proctoring?

This panel brought rich experience on the frontline of practice, business and academia:

Phillip Dawson is a Professor and the Associate Director of the Centre for Research in Assessment and Digital Learning, Deakin University. Phill researches assessment in higher education, focusing on feedback and cheating, predominantly in digital learning contexts. His 2021 book “Defending Assessment Security in a Digital World” explores how cheating is changing and what educators can do about it.

Jarrod Morgan is an inspiring entrepreneur, award-winning business leader, keynote speaker, and chief strategist for the world’s leading online testing company. Jarrod founded ProctorU in 2008, and in 2020 led the company through its merger and evolution into Meazure Learning. In his role as chief strategy officer, he is a frequent speaker for the Online Learning Consortium (OLC), the Association of Test Publishers (ATP), Educause, and many others. He has appeared on PBS and the Today Show, and has been covered by the Wall Street Journal, The New York Times, and is a columnist with Fast Company through their Executive Board program.

Jeannie Paterson is Professor of Law and Co-Director of the Centre for AI and Digital Ethics, University of Melbourne. She teaches and researches in the fields of consumer protection law, consumer credit and banking law, and AI and the law. Jeannie’s research covers three interrelated themes: The relationship between moral norms, ethical standards and law; Protection for consumers experiencing vulnerability; Regulatory design for emerging technologies that are fair, safe, reliable and accountable. She recently co-authored “Good Proctor or “Big Brother”? Ethics of Online Exam Supervision Technologies”.

Lesley Sefcik is a Senior Lecturer and Academic Integrity Advisor at Curtin University. She provides university-wide teaching, advice, and academic research within the field of academic integrity. She is a Homeward Bound Fellow and a Senior Fellow of the Higher Education Academy. Dr. Sefcik’s professional background is situated in Assessment and Quality Learning within the domain of Learning and Teaching. Current projects include the development, implementation and management of remote invigilation for online assessment, and academic integrity related programs for students and staff at Curtin. She co-authored “An examination of student user experience (UX) and perceptions of remote invigilation during online assessment”.

(Chair) Simon Buckingham Shum is Professor of Learning Informatics and Director of the Connected Intelligence Centre, University of Technology Sydney, where his team researches, deploys and evaluates Learning Analytics/AI-enabled ed-tech tools. He has helped to develop Learning Analytics as an academic field for the last decade, and has served two terms as SoLAR Vice-President. His background in ergonomics and human-computer interaction always draws his attention to how the human and technical must be co-designed to work together to create sustainable work practices. He recently coordinated the UTS “EdTech Ethics” Deliberative Democracy Consultation in which online exam proctoring was an example examined by students and staff.

Further resources shared during the webinar:

Explainable AI in Education

A new piece (open access), with thanks to Hassan Khosravi for coordinating this…

Hassan Khosravi, Simon Buckingham Shum, Guanliang Chen, Cristina Conati, Yi-Shan Tsai, Judy Kay, Simon Knight, Roberto Martinez-Maldonado, Shazia Sadiq, Dragan Gašević (2022). Explainable Artificial Intelligence in EducationComputers and Education: Artificial Intelligence, Vol. 3, 2022, 100074, ISSN 2666-920X. DOI: https://doi.org/10.1016/j.caeai.2022.100074

There are emerging concerns about the Fairness, Accountability, Transparency, and Ethics (FATE) of educational interventions supported by the use of Artificial Intelligence (AI) algorithms. One of the emerging methods for increasing trust in AI systems is to use eXplainable AI (XAI), which promotes the use of methods that produce transparent explanations and reasons for decisions AI systems make. Considering the existing literature on XAI, this paper argues that XAI in education has commonalities with the broader use of AI but also has distinctive needs. Accordingly, we first present a framework, referred to as XAI-ED, that considers six key aspects in relation to explainability for studying, designing and developing educational AI tools. These key aspects focus on the stakeholders, benefits, approaches for presenting explanations, widely used classes of AI models, human-centred designs of the AI interfaces and potential pitfalls of providing explanations within education. We then present four comprehensive case studies that illustrate the application of XAI-ED in four different educational AI tools. The paper concludes by discussing opportunities, challenges and future research needs for the effective incorporation of XAI in education.

What capabilities do learners need for an AI world?

I took part in an enjoyable and stimulating exercise led by my colleague ‘up the road’, the fabulous Lina Markauskaite, in which she orchestrated a “polylogue” among authors on how we envisioned the capabilities that learners will increasingly need in an AI-infused society. We each responded independently to a common set of prompt questions, and then began to comment on each others’ work, moving to a discussion and synthesis. This produces a different kind of article. See what you think (open access)…

L. Markauskaite, R. Marrone, O. Poquet, S. Knight, R. Martinez-Maldonado, S. Howard, J. Tondeur, M. De Laat, S. Buckingham Shum, D. Gašević, and G. Siemens (2022), Rethinking the entwinement between artificial intelligence and human learning: What capabilities do learners need for a world with AI? Computers and Education: Artificial Intelligence, Vol.3, 100056. https://doi.org/10.1016/j.caeai.2022.100056

The proliferation of AI in many aspects of human life—from personal leisure, to collaborative professional work, to global policy decisions—poses a sharp question about how to prepare people for an interconnected, fast-changing world which is increasingly becoming saturated with technological devices and agentic machines. What kinds of capabilities do people need in a world infused with AI? How can we conceptualise these capabilities? How can we help learners develop them? How can we empirically study and assess their development? With this paper, we open the discussion by adopting a dialogical knowledge-making approach. Our team of 11 co-authors participated in an orchestrated written discussion. Engaging in a semi-independent and semi-joint written polylogue, we assembled a pool of ideas of what these capabilities are and how learners could be helped to develop them. Simultaneously, we discussed conceptual and methodological ideas that would enable us to test and refine our hypothetical views. In synthesising these ideas, we propose that there is a need to move beyond AI-centred views of capabilities and consider the ecology of technology, cognition, social interaction, and values.

Deliberative Democracy for EdTech Ethics

What principles should govern UTS’ use of analytics and artificial intelligence to improve teaching and learning for all, while minimising the possibility of harmful outcomes?

This was the challenge we set a team of 20 people – students, casual tutors and full-time academics. And 5 intensive workshops later, they had delivered their response! A draft set of ethical principles to govern the use of these fast-changing technologies in UTS. How did we manage this? Below is the executive summary from the report on the EdTech Ethics website.

Executive Summary

This report has been written to document a novel community consultation process, using the principles and methods of Deliberative Democracy to consult with the UTS community on the following brief:

What principles should govern UTS use of analytics and artificial intelligence to improve teaching and learning for all, while minimising the possibility of harmful outcomes?

We’re sharing this to assist colleagues in UTS and beyond who are seeking more participatory models for community deliberation, with (in this case) specific application to the responsible use of educational technology that is powered by analytics and artificial intelligence. This is not a research paper, seeking to argue conceptual or empirical contributions to academic fields, although research is underway analysing and evaluating this process. We do hope, however, that this represents an interesting and novel ‘data point’ that others will find useful.

Deliberative Democracy (DD) is a movement in response to the crisis in confidence in how typical democratic systems engage citizens in decision making. DD works by creating a Deliberative Mini-Public (DMP). DMPs can be convened at different scales (organisation; community; region; nation) and can take many forms.

A DMP of 20 was selected through stratified sampling from UTS students, casual tutors and academics, who engaged in a series of five online workshops over seven weeks, due to Covid-19 conditions. With little to no prior knowledge among most members, they learned about the topic, worked well together, and converged on a set of principles that they felt reflected their shared values. The university experts who were involved in the workshops recognised the quality of the progress made in such a short period. UTS now has a plausibly representative expression of the community’s values, interests and concerns, in response to the brief. The principles can be viewed in Appendix 1: Draft Ethics Principles.

The raison d’etre for the initiative is to build trust within the university that these technologies are being deployed responsibly. The DMP process delivered on its promise to build engagement and trust across diverse stakeholders. The recording of the final briefing (18 mins, below) conveys the passion and commitment that the DMP invested in the process and outcome, reinforced by the preliminary themes emerging from interviews with students, educators and senior leaders.

Deliberative Democracy, even when conducted wholly online, would appear to offer educational institutions an approach to address the urgent need for meaningful student/staff consultation on the ethical implications of introducing Learning Analytics and Artificial Intelligence into teaching and learning. The implementation process is now beginning, which we will be studying with equal interest.