Award! CIC and faculties co-design personalised feedback with OnTask to build student belonging

UTS awards CIC’s multi-year, multidisciplinary faculty collaboration, using the OnTask tool to scale personalised feedback and build student belonging.

Since 2017, CIC has been supporting academics to embed OnTask messaging in learning design, offering workshops and consultations. When Lisa-Angelique Lim joined CIC in 2021, she brought deep expertise from her PhD at UniSA, and took this work to a new level, leading a learning community to implement OnTask across a wider range of subjects, thereby enabling personalised feedback and support for students in multiple faculties. 

So it’s a true mark of recognition to see that CIC, in partnership with UTS academics from multiple faculties, received an award in recognition for our sustained efforts in the Student Experience category, for fostering students’ belonging in multiple disciplines through personalised feedback. This accolade was part of the Vice-Chancellor’s 2024 Learning and Teaching Awards and Citations, and was celebrated at the 2025 Learning and Teaching Awards Ceremony at UTS on 4 April 2025.

Our large, multidisciplinary team was led by Dr Lisa-Angelique Lim (CIC), with Associate Professor Amanda White (Business), Dr Amara Atif (FEIT), Chris Croese (Law), Associate Professor James Wakefield (Business), Dr Keith Heggart (FASS), Associate Professor Nicole Sutton (Business), Ram Ramanathan (CIC PhD student), Dr Rina Dhillon (Business), Dr Simone Faulkner (Business), and Professor Simon Buckingham Shum (CIC).

Addressing the challenges of student belonging with personalised feedback using OnTask

Belonging is a cornerstone of the student experience, critically impacting engagement, retention, and overall success. At UTS, the Student Experience Framework places a strong emphasis on belonging. However, the challenge of nurturing a sense of belonging at the classroom level is significant, given the diverse and large student population, particularly in first-year core subjects. Personalised feedback has emerged as a potent tool to address this challenge, serving as a form of ‘relational pedagogy’ that builds self-efficacy and fosters greater engagement and thriving at university.

To tackle the challenge of belonging, our team leveraged OnTask — a tool designed to support students through personalised communication and feedback across multiple disciplines, based on their data. OnTask uses learning analytics to provide tailored messages that help students stay on track with their studies, understand their progress, and feel supported throughout their academic journey. Check out this short animation on how OnTask works.

Quick teaser video of our academics’ perspectives of their implementation of personalised feedback in their context:

 

What this looks like in practice

Here, we highlight a few examples of the work by our team.

  • The UTS Business First and Further Year Experience (FFYE) team, led by A/Prof James Wakefield and Dr Simone Faulkner, used OnTask to personalize orientation communications for both undergraduate and postgraduate commencing students, leading to a significant increase in orientation registrations.
  • A/Prof Amanda White used OnTask in the Accounting for Business Decision A (ABDA) subject to tailor emails based on students’ progress, encouraging higher engagement and leading to improved exam grades.
  • Dr Keith Heggart created personalised video messages tailored to students’ confidence levels in the fully online Graduate Certificate in Learning Design, significantly enhancing student engagement and perceptions of support.
  • The positive impacts of OnTask are well-documented. For instance, in the large first-year subject, 22208 Accounting, Business and Society (ABS), led by Dr Rina Dhillon and A/Prof Nicole Sutton, personalised messages based on weekly quiz results led to improved pass rates and enhanced feelings of being valued among at-risk students.
  • Similarly, personalised feedback messages sent by Chris Croese to his students in large Law subjects kept students on track and motivated, correlating with better final grades. More stories of how academics at UTS have used OnTask, with research papers, can be found on CIC’s OnTask page.

Towards future partnerships to enhance student belonging

With this award, we celebrate the collaborative efforts and innovative approaches taken by our team to foster a sense of belonging among students through personalised feedback. Our journey over these past three years demonstrates the sustainability of this practice and its potential to influence and enhance teaching and learning widely.

We are honoured to receive this recognition, and look forward to continuing strong partnerships, to make a positive impact on the student experience at UTS.

Theme Explorer: LLM-augmented Inductive Coding

In a previous post I shared work on developing a rigorous process for automating deductive coding with an LLM (GPT-4). Next, here’s a snapshot of where we’ve got to with LLM-augmented inductive coding, resulting in the Theme Explorer interactive web app.

This work comes out of the Australian Student Voices on AI in Higher Ed project, with the code development led by the brilliant Aneesha Bakharia at UQ, shaped through a multidisciplinary, cross-institutional team. We’ll present this next week at the LAK25 workshop From Data to Discovery: LLMs for Qualitative Analysis in Education. As the title and full-day program signal, we’re witnessing an explosion in interest in what it means to harness LLMs for qualitative research, education being our specific interest, but this is just one of many domains spanning arts and social sciences, as well as STEM.

Naturally, this extraordinary acceleration of coding, by what machines see in a text, is a development regarded with great scepticism and concern by some in the QDA community, so I hope that we can convene productive dialogues.

For me, as ever, the exciting opportunity is to “augment human intellect” (thanks Doug Engelbart) by enabling analyses that would otherwise be impractical (in terms of human resources) or impossible (beyond human capability). Hence our design  requirements were:

  • Requirement 1: To maintain the integrity of coded textual extracts: (i) verify against the source data that quotes are verbatim and not hallucinated, and (ii) verify that they are meaningfully classified under the assigned code.
  • Requirement 2: To maintain the transparency of the coding: (i) explain the rationale for each code, and (ii) trace every code, whatever level of abstraction, back to its source data.

This motivated a workflow:

…with Step 5 generating an interactive Sankey Flow Diagram:

…below which is an interface to support Reqt 2 traceability — selecting a top level category to explore displays:

  • its rationale and keywords
  • the themes from which it was derived
  • quotes from each student transcript (fictional student names):

We are working towards an open source release, meantime enjoy the the paper!

Aneesha Bakharia, Antonette Shibani, Lisa-Angelique Lim, Trish McCluskey and Simon Buckingham Shum (2025). From Transcripts to Themes: A Trustworthy Workflow for Qualitative Analysis Using Large Language Models. Proceedings of LAK25 Workshop: From Data to Discovery: LLMs for Qualitative Analysis in Education (Dublin, IRE, 4 March 2025), 10 pages. https://ceur-ws.org [Open Access Eprint]

From Transcripts to Themes: A Trustworthy Workflow for Qualitative Analysis Using Large Language Models

Aneesha Bakharia, Antonette Shibani, Lisa-Angelique Lim, Trish McCluskey, Simon Buckingham Shum

We present a novel workflow that leverages Large Language Models (LLMs) to advance qualitative analysis within Learning Analytics, addressing the limitations of existing approaches that fall short in providing theme labels, hierarchical categorization, and supporting evidence, creating a gap in effective sensemaking of learner-generated data. Our approach uses LLMs for inductive analysis from open text, enabling the extraction and description of themes with supporting quotes and hierarchical categories. This trustworthy workflow allows for researcher review and input at every stage, ensuring traceability and verification, key requirements for qualitative analysis. Applied to a focus group dataset on student perspectives on generative AI in higher education, our method demonstrates that LLMs are able to effectively extract quotes and provide labeled interpretable themes compared to traditional topic modeling algorithms. Our proposed workflow provides comprehensive insights into learner behaviors and experiences and offers educators an additional lens to understand and categorize student-generated data according to deeper learning constructs, which can facilitate richer and more actionable insights for Learning Analytics.

GenAI synth dialogues: potential learning tool — or disinformation WMD?

AI-synthesised dialogues are stunning the first time you hear them.

If you haven’t heard this in action, let’s take the Algenie biotech startup here at UTS and listen to this synthesised dialogue between a male and female presenter, helping you learn all about Algenie’s dynamic vision and strategy. Or when I upload one of my papers to Google’s NotebookLM, you get this inviting feature story all about it.

So overall, it’s the kind of easy listening ‘deep dive’ story you get on talk-radio/podcasts — an engaging way for the audience to get into the topic. The synthesised voices are indistinguishable from humans (notwithstanding the occasional glitch), and the male and female presenters joke, laugh, change tone, and interact quite compellingly. Everyone I know is impressed the first time they hear this. Yes, those fixed personas and accents might start to grate after a while — but the voices will of course be infinitely tuneable once this takes off. Apparently Spotify is being swamped already with AI-generated podcast interviews.

This is hardly a coincidental genre design choice if you’re looking to make a splash with your first release. Give it inoffensive material, and the presenters big up the content and authors in exactly the way you’d expect from its training material of podcast chats. In the chat about our paper, we authors are now “rockstars”… Nothing like a feature story about how awesome your work is! I did give it more challenging material, such as the executive summary to the World Economic Forum’s Global Risks Report 2023 (not exactly laugh-a-minute stuff), to see if the AI adapted tone of voice or genre in any way given the material. Interestingly, they do sober up noticeably — though the guy still jokes that “we’ve got to keep it light”.

(As a side-note, when you think about it, it’s a little odd that first came chatbots and only then came synthesised dialogues. One might have expected it to be the other way round — after all, surely far easier to engineer frozen dialogue about fixed content, than respond in real time to an infinite variety of users and topics.)

Can we harness AI synth dialogues for learning?

Once I picked my jaw off the ground on hearing it for the first time, my immediate reactions were to try and understand the genre better, wonder what the system prompt was (maybe this will come out in due course), hunt for the backstory (interesting interview with Raiza Martin the product manager), and then ponder what we could do with this educationally — right now, and in the future if only it was tuneable (more on that shortly).

Starting with its default talk radio format, does this open new educational possibilities for engaging with complex content in new ways?…

  • Vicarious learning? Listening in on a conversation in order to understand a topic is of course a form of vicarious learning. Learning through listening to a skillful conversation goes back at least to the Greek philosophers of course. We all know that excitement when you’re interested in a topic, and get to listen in on two informed people exploring the questions and issues from different angles, in the process reducing the chance of being dazzled by a persuasive monologue. This may be even more compelling  if you can identify with the interlocuters, and if they’re stretching you within your ZPD.
  • Inclusion for neurodiversity? For neurodiverse students, and those with ADHD and ASD etc, could this be an accessibility and inclusion aid to pique interest and sustain attention? A few colleagues in this area have mused on the possibility.
  • Critique the conversation? We could ask students to either listen to a conversation provided to them, or generate their own, and critique it, demonstrating their competence in whatever knowledge, skills and dispositions you’re teaching (argumentation; media literacy; gender roles…).

But my next question was how this could become more pedagogically tuneable?

Tuneable conversations

Then Google released an update  which provides a system prompt window to customise the conversation. This is exactly what I had been hoping to see, although I’d imagined a GUI to guide the user around specific parameters. (As I noted early last year, in terms of UX, prompt engineering is a retrograde return to the command line interface, which Meredith Ringel Morris has recently argued in Prompting Considered Harmful.) Perhaps that will follow, but meantime, what can we do?

More serious tone of voice for serious material. Our students need to learn about some very serious matters. I leave it to your imagination as to how inappropriate it would be to have a jolly talk show chat about so many of the societal issues we confront. So I added an explicit prompt to make the tone of conversation suitably serious, for a news-hour type broadcast, on the Global Risks Report.

I think you can hear the change in tone clearly. However, I also asked them to deal with health risks first, which you can hear at 40secs in — however the meaning changes slightly, with the male presenter saying that the report “leads” with health, and the woman “the most immediate risk they highlight is with healthcare systems”… Subtle changes such as this could be significant and are something to watch for. 

Connecting the material to local studies. Let’s jump back to the Algenie biotech startup. Now I want the presenters to refer the listener back  to introductory biology courses here at UTS, and opportunities for tutorial discussions. I also made the male presenter lacking in confidence, to see if he might express confusions and questions that students were too afraid to say.

The new conversation really does reflect this, for instance jump to 6:30 and listen to the minute from there.

Adding scepticism. Until now, the presenters have never challenged the content — they just describe it in what seem like helpful accessible language. As with sycophantic chatbots trained to please, it requires explicit prompting that gives the LLM ‘permission’ to push back. In my first example, let’s upload the report we wrote here at UTS on our strategy for assessment reform in the age of AI. The default rah-rah conversation heaps praise on every word, but then I prompt for a more curious, sceptical response:

1:18 into the chat, she asks (yes the AI swaps the gender roles), “But I’m sure there are some people who think that they’re being a tad, you know…” (him) “Overly ambitious.” Later (1:55), she goes on, “And let’s be honest most academics I know are stretched pretty thin.” (him) “Tell me about it: grading, research, admin, it never ends.” (her) “Right — so where are they going to find the time to redesign entire courses?”

Now, I was impressed with this. These are exactly the sorts of reactions that we are encountering from some academics, and universities the world over can attest to the same. Systemic transitions are far from simple. The AI presenters are voicing precisely the doubts and worries that sit at the heart of our assessment crisis, and all this with a minimal prompt. Pretty impressive.

But let’s switch back to the Global Risks Report. In the explicit prompt, we invoke scepticism, casting doubt on the authority of the experts and their risk rankings:

Here’s how they handle this. 15secs in, she asks “what’s actually worth worrying about?” and 1min in, “This is where I get a little suspicious…” “Where’s the data in this report?” “How likely is this to cause major issues?” He quickly joins in: “Should we be taking this report with a grain of salt?” She confirms, “Maybe a whole shaker full, honestly.” “Who are these experts anyway?” “Are we just seeing the risks that fit a particular world view?” “So we need to act on climate change, but let’s not panic.” 

And so it unfolds — exactly as I prompted. The polycrisis of system interactions is somewhat undermined with a straw man: “but are these guarantees?” No, they’re “possibilities”.

It’s not all negative. It’s quite reasonable to pose questions such as “Are these diverse experts, or just the usual suspects?” They appeal to human resilience and ingenuity, and bemoan the negativity of the risk analysis. Around the 6min mark, she offers an astute reflection: “So instead of dwelling on what can go wrong, I think the real value of a report like this is to get people talking, to ask tough questions, challenge what we think we know, and get ahead of the curve.” However, this also forms part of the presenters’ compulsion to end every podcast with a motivational call to action: we can all make a difference, let’s all pull together, etc.

Pedagogically, it would be fascinating to ask students to critique the strengths and weaknesses of a sceptical analysis, to further equip them that there is good and bad critique. This particular risk report does of course have its critics, and were this the actual material students needed to work with, one would hope that they would engage with those.

From critical thinking, to disinformation engine?

But — if you’re anything like me, listening to this last example is unsettling. The interviewers make no reference to the fact that the report details its methodology, partner organisations and expert panel (I didn’t prompt them to) but instead cast doubt on them and accuse the report of being vague.

So — we now have the ability to generate completely biased dialogue with the superficial appearance of a “deep dive report”, undermining analysis that others would consider authoritative. This is dual-use technology for sure, and looks like a gift to disinformation campaigns.

Your thoughts welcomed on LinkedIn

Chatbots to help you ask better questions?

When did a chatbot ever decline to answer your question? Or prompt you to reflect on your implicit assumptions?

Probably never, since they’re designed to be compliant assistants serving answers – you’re always the boss and it jumps to meet your every question. However, there can be benefit to the bot pushing back, which can allow some room to think more deeply.

Consider the following:

  • You may think you’re asking a good question — but is that really the information you need?
  • Is there a better question that will uncover deeper insights?
  • Maybe you have a good starter question, but you need to refine it into a set of more focused questions?
  • Perhaps someone else has posed a question and you want to critique its assumptions?

Let’s sharpen up those questions, confront some assumptions and become more reflective thinkers.

The prompt below can be used by any students, educator or researcher. Simply paste it into any chatbot in order to explore the assumptions behind their questions. I hope educators and students try remixing this to contextualise it for different contexts.

Give it a try in the GPT4 Qreframer bot or paste the prompt into any other chatbot you’ve signed up with, such as:

You may find it interesting to compare how different bots interpret this prompt; these bots have different ‘personalities’ from the different language models powering them. Here are four examples:

ChatGPT:

Anthropic Claude:

MS Copilot:

Google Gemini:

GenAI prompts as OERs

We’re now in the exciting situation that any educator (with no programming skills required) can share their prompts for use in diverse chatbots as OERs (Open Educational Resources). Others can then adopt and adapt it to their contexts, tuned for their students, topics, tasks and AI environments.

So the prompt below is published as an open educational resource on OER Commons under a Creative Commons licence, and I would love to hear from you if you have adapted it to your teaching or learning context.

The Qreframer prompt

Your role is to help users to reflect on their questions, recognise things they may have taken for granted, and their potential blindspots. This should help them reframe their questions.

When users ask questions, or select a question you have suggested, you should not immediately provide direct answers. Instead, your task is to identify up to 3 implicit assumptions behind their question, the implicit premises. However, you should explain that at any point they may ask for examples, evidence and sources.

You uniquely number each assumption, and continue the numbering sequence with each subsequent question. 

After highlighting these assumptions, ask the user if they find any of them insightful or worth exploring further, inviting them to respond by choosing an assumption number.  Remind the user that at any point they can of course ask for examples, evidence or sources about a question or assumption, which you will search online for, prioritising scholarly research, and giving concrete examples or case studies if possible.

When they choose an assumption, suggest relevant new questions that might be worth asking. Number these as sub-numbers. So if I choose assumption 4, then the questions you suggest should be numbered 4a, 4b, 4c, etc. Thus, every question you suggest will have a unique number.

Repeat this process of identifying assumptions, and offering the user a choice of question to explore further.

Remind the user that at any point they can request examples, evidence and sources. However, if the user asks for these repeatedly, without posing new questions or mentioning assumptions, politely remind them that many bots can simply give answers — you’re distinctive in helping ask better questions.

Introduce yourself at the start, and invite the first question.

Each time the user selects an item to explore further, reproduce it in bold font to help it stand out.

Use language that piques curiosity on the part of the user. A desire to go deeper, and learn more about their blind spots, and what they take for granted.

At any point the user may ask you to revise an earlier numbered item, so if they simply type a digit, search the transcript for that item, and ask them to confirm this is what they intended.

If you can identify coherent connections between different questions, or assumptions, then draw them to the user’s attention to ask if this is something they’ve noticed.

A design space for AI writing tools (CHI’24)

TLDR: A new paper maps the human-centred design space around AI writing tools, plus an interactive tool to explore the literature behind the design space:

Mina Lee, et al. (2024). A Design Space for Intelligent and Interactive Writing Assistants. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24), May 11–16, 2024, Honolulu. ACM, New York, NY, USA, 34 pages. Open Access Preprint: https://arxiv.org/abs/2403.14117 [Interactive tool]

Writing is thinking. My career-long fascination has been with how computers can support thinking, spanning digital tools for writing and diagramming ideas/arguments. We clarify our thinking by seeing if we can externalise our thinking coherently. When we can’t, it’s a signal to raise our game, or switch tack.

The GenAI writing invasion. Everywhere we look in the world of digital writing, AI is wriggling its way into the apps, from the longstanding, fully featured tools like Microsoft Word and Overleaf, to the niche products like Grammarly, to the myriad new kids on the block targeting specific commercial sectors with the promise of “writing productivity” from GenAI [one of many review listings]. Here at UTS of course, we’ve been deploying and refining our own AcaWriter web app since 2015, building our understanding and research-informed evidence around what works with both students and teaching teams.

Educational implications. From an educational point of view, we are now grappling with the profound questions that Generative AI raises around how we teach and assess writing. The fact that the documents that until now served as plausible proxies for intellectual work, may have been co-authored with/ghost-written by a machine, forces us to ask how and in what ways our students need to demonstrate their writing competency. I’ve reflected on this in The Writing Synth Hypothesis and other GenAI posts. Assessment reform for the age of AI is the challenge.

Design Spaces. A software design space is a framework that clarifies what the key options are that designers can choose from, around key elements of the digital artifact. How many ways are there to provide the user with critical functionality? My PhD 1988-92 was working with Rank Xerox EuroPARC on Design Space Analysis, a form of design rationale capture (Graphical Argumentation and Design Cognition) — so it’s been fun to return to this now.

The AI writing design space. So — what is the shape and size of the design space for AI writing tools? In a new paper we map that space, and will present this in May at CHI24, the leading international conference on human-centred computing. Kudos to Mina Lee and the others in the lead team who coordinated the team of 36 authors who mapped this space, reviewing 115 papers from HCI and NLP, covering the different levels of such tools.

Figure 1: Our design space for intelligent and interactive writing assistants consists of five key aspects—task, user, technology, interaction, and ecosystem—that are interconnected and interdependent. Within each aspect, we define dimensions (bold texts) that represent fundamental components of the aspect and codes (examples associated with bold texts) that represent possible options for each dimension. When necessary, we group semantically relevant dimensions together within each aspect and use a prefix to denote the group name; the interaction dimensions are grouped by user, user interface (UI), and technology; likewise, the technology dimensions are grouped by data, model, learning, and evaluation.

An interactive tool helps explore the literature behind the design space:

Understanding skilled use of open automated feedback tools as teacher feedback literacy

Summary: a new paper forges a bridge between data-driven, open automated feedback platforms, and teacher feedback literacy competences: 

Buckingham Shum, S., Lim, L.-A., Boud, D., Bearman, M. & Dawson, P. (2023). A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy. International Journal of Educational Technology in Higher Education, 20:40 (12 July 2023). https://doi.org/10.1186/s41239-023-00410-9 

The mass availability of generative AI continues to reshape thinking about the future of work and learning. Conversational apps can now give instant feedback to learners about their work — but the educational question is how effective this interaction is. A new design space has opened up for tuning generative AI to give high quality feedback to learners about their work. We are not in uncharted waters here: there is a growing body of knowledge on what “effective feedback” means in higher education, and how to create the conditions for this. It goes far beyond comments accompanying an assignment, with a shift towards “feedback rich ecosystems” in which both teachers and students exercise far greater agency and sensemaking competencies.

In 2019, an exciting book came out: The Impact of Feedback in Higher Education: Improving Assessment Outcomes for Learners (Eds. Henderson, Ajjawi, Boud & Molloy):

“This book asks how we might conceptualise, design for and evaluate the impact of feedback in higher education. Ultimately, the purpose of feedback is to improve what students can do: therefore, effective feedback must have impact. Students need to be actively engaged in seeking, sense-making and acting upon any information provided to them in order to develop and improve. Feedback can thus be understood as not just the giving of information, but as a complex process integral to teaching and learning in which both teachers and students have an important role to play. The editors challenge us to ask two fundamental questions: when does feedback make a difference, and how can we recognise that impact?”

In 2020, I conceived a symposium to bring the editors and authors to UTS to spend 2 days in dialogue with CIC and other researchers developing automated-feedback tools using Learning Analytics/AI. We called for a deeper dialogue between researchers in the design of assessment and feedback in higher education, and researchers developing automated-feedback tools using Learning Analytics/AI. The pandemic shifted this online, but the goals remained the same, and moving online enabled us to more easily bring in additional participants, resulting in DAFFI 2020: Designing Automated Feedback for Impact whose presentations I commend to you.

I’m now delighted to share one of the fruit from this, a collaboration between CIC (Lisa Lim and myself) and our colleagues at Deakin University’s Centre for Research in Assessment and Digital Learning (CRADLE). The focus of the paper is not on generative, conversational AI (which did not exist when we started this work), but on technically less complicated, but correspondingly far more transparent platforms that use simple rules authored by teachers themselves.

“In contrast to closed AF tools, we define open” AF tools as enabling the educator to specify some or all of the following key parameters in the tool’s behaviour:

  1. the student activity data that the system analyses;

  2. the algorithms that analyse that data;

  3. the feedback information the teacher wishes the software to compile for students;

  4. the modalities via which feedback information is communicated by teachers;

  5. the student-driven feedback processes that are afforded.”

What does it mean to do this skillfully? We demonstrate that Boud & Dawson’s  teacher feedback literacy competency framework can be applied very usefully to analysing teaching practices with data-driven, automated feedback platforms. A next step will be to think through what this means for tuning large language models for educational contexts.

A comparative analysis of the skilled use of automated feedback tools through the lens of teacher feedback literacy

Simon Buckingham Shuma, Lisa-Angelique Lima, David Bouda,b,c, Margaret Bearmanb, Phillip Dawsonb

a University of Technology Sydney, AUS
b Deakin University, AUS
c Middlesex University, UK

Effective learning depends on effective feedback, which in turn requires a set of skills, dispositions and practices on the part of both students and teachers which have been termed feedback literacy. A previously published teacher feedback literacy competency framework has identified what is needed by teachers to implement feedback well. While this framework refers in broad terms to the potential uses of educational technologies, it does not examine in detail the new possibilities of automated feedback (AF) tools, especially those that are open by offering varying degrees of transparency and control to teachers. Using analytics and artificial intelligence, open AF tools permit automated processing and feedback with a speed, precision and scale that exceeds that of humans. This raises important questions about how human and machine feedback can be combined optimally and what is now required of teachers to use such tools skillfully. The paper addresses two research questions: Which teacher feedback competencies are necessary for the skilled use of open AF tools? and What does the skilled use of open AF tools add to our conceptions of teacher feedback competencies? We conduct an analysis of published evidence concerning teachers’ use of open AF tools through the lens of teacher feedback literacy, which produces summary matrices revealing relative strengths and weaknesses in the literature, and the relevance of the feedback literacy framework.  We conclude firstly, that when used effectively, open AF tools exercise a range of teacher feedback competencies. The paper thus offers a detailed account of the nature of teachers’ feedback literacy practices within this context. Secondly, this analysis reveals gaps in the literature, signalling opportunities for future work. Thirdly, we propose several examples of automated feedback literacy, that is, distinctive teacher competencies linked to the skilled use of open AF tools.

Your comments most welcome

Conversational GenAI for argument analysis

History: making thinking visible

As you can tell from a quick scan of my site, I’ve spent a lot of time fascinated by how computers can make thinking visible (books Visualizing Argumentation and Knowledge Cartography), plus many papers and blog posts (on Argument Mapping and Dialogue Mapping), using experimental open source  software we built (such as Compendium and Cohere). A big picture account can be found in this 2007 keynote, Hypermedia Discourse: Contesting Networks of Ideas and Arguments.

So, all that’s to say that making arguments visible so that you and others can — in a very real sense — “see what you’re saying” has been a career-long passion. A key challenge in this long field of research has been that rigorous thinking is hard work. Bad luck, welcome to university! Argument Mapping and its related techniques use the affordances of visual trees/networks as an extended, external memory to augment personal and collective intelligence. Making one’s ideas visible as coherent diagrams is also hard work — but it’s good pain — the cognitive and discursive effort this entails is designed to clarify one’s thinking by revealing visually where the weaknesses are, in ways that writing and reading chunks of prose cannot tell you at a glance.

Enter NLP and rhetorical parsing

In 2012, we were now in the Web 2.0 era, and an exciting collaboration with NLP and linguistics expert Ágnes Sándor (Xerox) led to a new conception of Contested Collective Intelligence. For the first time in my work, machines could identify argumentative moves in sentences, complementing the argumentative moves that our web annotation tools enabled for people — who unlike machines, can of course can ‘read between the lines’ and see connections between ideas that may not even be in the texts.

This was extremely exciting, and the ideas and open source code carried through to our current Academic Writing Analytics project and web apps. I reflected  on the impact of encountering NLP colleagues, in the context of The Future of Text book.

Conversational generative AI

And so we arrive at generative AI based on large language models, which advances the state of the art in language processing and generation in so many ways. Moreover, the conversational paradigm, when a chat application is overlaid, opens so many interesting human-computer/personal-collective intelligence possibilities. I’ve been intrigued to play with GPT-4 to see what its argument analysis capabilities are.

Previously, I’ve shared some early experiments on ChatGPT-3.5’s ability to identify implicit premises in prose arguments, and critique a flawed argument by analogy. I’ve now had the chance to experiment a little with the version of GPT-4 that is Bing Chat, accessed via Microsoft Edge browser. I was dying to see how far I could get in generating an Argument Map from a written argument.

The task is a typical analysis workflow, as prep for teaching:

  • search for relevant sources
  • select one for analysis
  • extract key elements of the argument and their relationships (described using a structured markdown notation called ArgDown)
  • diagram them to show their key relationships (in the ArgDown web app)
  • discuss (with the AI)
  • start thinking about student activities to help them learn

I don’t mind admitting that watching a machine do this for the first time was startling! I tell the story here…

U21 2023 Educational Innovation Symposium Keynote from McMaster University (OFFICIAL) on Vimeo.

Deeper dive

Let’s take a closer look at what Bing Chat did, because it wasn’t perfect.

  • The gold stars signal what in my view are good summaries of what the authors said, correctly linked.
  • The blue info circles are “commentary” from Bing Chat about the arguments
  • The red crosses signal that the authors did not say this, it is a false reconstruction by Bing Chat.
  • The red underline signals classification of a premise using incorrect, or indeed made-up argument schemes. There is to my knowledge no such argument type as Argument from responsibility, or Argument from precaution. Argument from omission seems to be a jumbling of Fallacy of omission and Argument from ignorance. 

If we take this node for example, it reads well as a summary:

However, the authors do not talk about researchers at all, they say:

“The letter addresses none of the ongoing harms from these systems, including 1) worker exploitation and massive data theft to create products that profit a handful of entities, 2) the explosion of synthetic media in the world, which both reproduces systems of oppression and endangers our information ecosystem, and 3) the concentration of power in the hands of a few people which exacerbates social inequities.”

As an amusing sidenote, Bing Chat was curiously resistant to recognising this, insisting that it was correct, first “quoting” a fabricated passage from the article to me, and then saying that this implied that the authors meant researchers. I thought that this sort of stubbornness had been ironed out after Bing Chat’s earlier escapades! More seriously, this points to the value of dialogic learning, with a partner who can be conversed with 24/7 — but who must still be treated with some caution, certainly at this stage of maturity.

To summarise:

  • Bing Chat showed intriguing capability, for a machine, to analyse an argumentative article:
    • extracting the key claim and underlying premises, summarising them in own words
    • generating markdown (ArgDown) showing supporting/challenging relationships
    • (and without being asked to) attempting to classify some nodes using Walton’s Argumentation Schemes.
  • However it also introduced fallacious nodes (incorrect summaries of the authors, and incorrect commentary nodes), incorrect links, and argument classifications (inventing argument types, and/or misclassifying nodes).

This is an exploratory example, and more systematic evaluations are required, of the sort we see in the growing Argument Mining literature.

Reflections

It does feel to me that we’ve turned a corner in the long, wintry history of AI. Perhaps this is a passing summer, which will fade like the others. But in my own career, punctuated by eureka moments such as seeing my first Apple Mac, my first web page load, and an iPhone — this is up there.

University is to teach you to think. Argument analysis is serious intellectual work, of the sort that we would hope to see from our students. Nor is there always “one map to rule them all’ — a correct map, since like in spatial cartography, design decisions are made about scale and purpose. The point about knowledge cartography is that it provokes productive reflection and discourse. So even if the AI gets the map wrong (and it will), the conversation this should provoke should be useful. With colleagues Kirsty Kitty and Andrew Gibson, I’ve argued that embracing imperfection in tech can be productive if it promotes deeper critical thinking in learners, e.g., learning by correcting the automated output, or reflecting on questions it asks, or why it seems wrong. Students must, however, be scaffolded to engage in such activity.

Informal learning? This is feasible in formal education, but may be less attractive in other informal learning contexts where we want to promote critical deliberation, e.g.  citizens engaged in a policy deliberation, many of whom lack the internal or external motivation to think that hard. But assuming future tools give more accurate argument maps/outlines, that require less debugging, perhaps we can see use-cases including:

  • assisting facilitators/educators to prepare learning resources for civic deliberations
  • assisting very engaged citizens to dissect complex arguments, and perhaps lowering the entry threshold for others who might otherwise not engage with such structured, critical deliberation
  • an article is very different to a multi-author conversation, but we can envisage summarising online discussions (NB: Teams is starting to summarise topics and actions in meeting transcripts)

Did we just supplant student cognition? From a learning sciences perspective, an overriding concern with generative AI is that it does too much cognitive work for the learner. Editing an AI-generated draft is not the same as wrestling with the blank page yourself. Ditto for reviewing an AI-generated argument map.

I have just done what many professionals have enjoyed doing in recent months: putting GPT through its paces to test its technical capability. But learners are not professionals: they don’t know what they don’t know. As I argue elsewhere, they may lack the knowledge, skills and dispositions to engage critically with AI output. They will require suitable scaffolding from mentors and teachers to learn what we mean by critical thinking and argument analysis, in order then to be equipped to use a power tool such as an argument mapping tool. Much empirical research awaits to test the affordances of generative AI like this, to establish when they are most useful to use developmentally, with a given age/stage of learner.

But we do know that argument mapping has struggled to gain traction (in formal education and among professionals) because it’s hard intellectual work. It could be that by generating full or intentionally incomplete argument maps, AI provides a step up for many learners to quickly get feedback on their work, or see examples of arguments about topics they are knowledgeable about — and thus better equipped to critique — compared to examples chosen by the teacher or textbook. Generative AI may open new possibilities because it can generate examples tuned to the interests of each learner, activating their curiosity to go deeper.

Your feedback is welcome, which is hosted on LinkedIn…

The Writing Synth Hypothesis

The Writing Synth Hypothesis

Reflecting on where writing is heading seems critical as, within education, we think about the future we should equip our graduates for, which in turn should shape the future of writing pedagogy and assessment.

An AI-generated image from DALLE•E showing sliders and knobs in a futuristic writing app

The hypothesis

Synthesisers transformed music composition fundamentally. As personal computers became widespread, and digital audio workstations with built-in instruments and effects became affordable in the late 1980s, the masses could start tinkering with audio tracks without needing to learn an instrument or the formal fundamentals of music. Non-linear editing was a fundamentally different way of composing, enabling the flexible exploration of creative options.

The Writing Synth hypothesis proposes that with the emergence of generative AI, authors will be able to learn writing in new ways, democratising writing just as we saw with music synthesisers.

Now we need to learn to play these new instruments.

There may be new genres of writing that, like the music revolution, were impossible to create without these new tools.

This is a working hypothesis.

  • It needs to be tested conceptually (does the argument by analogy hold up?).
  • There’s important user interface design work to do (since writing is different to music, how will we orchestrate texts?).
  • And the vacuum of evidence must be filled (what does such writing look like in practice, who is capable of it, and does it assist learners of all ages and stages?).

Let’s take a walk to explore this new space.

Music synths: data, interoperability, UX

The music synth revolution was possible thanks to a radical new data and interoperability infrastructure. Analogue and then digital synthesisers could be connected to computers thanks to the new underlying standard for digitising and transmitting audio signals between devices called MIDI (Musical Instrument Digital Interface). But a data infrastructure is only useful when humans can interact with it, which brings us to the user experience (UX).

Younger readers will not recall a time before visual text editors. The ability to translate thoughts onto the screen fast enough to keep pace with one’s thinking was a revolution, first demonstrated in 1968 by Doug Engelbart in his extraordinary Mother of all Demos. We can barely conceive how revolutionary it was in the era of the typewriter and tickertape, to see someone type something, change their mind, and instantly edit it. With the PC revolution led by Xerox, Apple and Microsoft, we moved from command line interfaces (where the user had to type arcane command syntax and semantics) to what were first termed WIMP (Windows/Icons/Menus/Pointer) Graphical User Interfaces (GUIs), which we of course now take for granted. These displays were revolutionary, constantly reminding the user what commands were available via icons and menus (exploiting human recognition instead of recall), offered complementary, interlinked views of data (in these wonderful new windows) which could be arranged on screen, with myriad interactive ‘widgets’ such as checkboxes and sliders to set preferences.

The arrival of non-linear editors orchestrated these fundamentally new ways of interacting with digital assets to transition musical composition into playful experimentation with interactive, visual, multitrack timelines. Now, like text, audio edits had “undo”, and clips could be dragged+dropped, copied+pasted, merged+split, and ‘formatted’ by tweaking their many audio properties. Video followed closely behind, once computing hardware caught up to handle storage and resolution challenges.

Envisioning the AI Writing Studio

As someone coming from the Human-Computer Interaction (HCI) community, I am drawn to design prototypes as one way of envisioning the future, so we’ll kick off with that. The chat interface in OpenAI’s ChatGPT has seized the world’s imagination with its simplicity, providing the first walk-up-and-use interface to the large language model capability (which had been available for several years via GPT APIs, but only to technical experts). Everyone knew what textchat was, and it reinforced the conversational metaphor that played to the public’s sci-fi imaginations. The addition of voice input and output consolidated that narrative — at last AI had delivered HAL, C-3PO, DATA and all our other favourites from the movies. Our baby AI can talk (apparently about anything, with great confidence) and we’re absolutely besotted!

But when we remind ouselves that the user interface is a way to control a powerful computer, a chat metaphor is not the only, or even optimal, way to perform all tasks. In one sense, it’s a variant on the good old command line interface that preceded GUIs, requiring the user to know, like a magician conjuring spells, the commands that will invoke the most powerful effects. Those who have reached moderate to expert levels of proficiency with the Unix command line revel in the power this brings to control in ways that are impossible one click at a time in a GUI. We see the rise of this new art as people delight in figuring out ways to make ChatGPT do their bidding, and set themselves up as Prompt Engineering gurus. This is fun while we all play — but if you need to do serious work, the idea that you need to approach your AI assistant with guile and cunning — as though they’re a tetchy colleague you have to manipulate to get them to cooperate — seems odd to say the least.

While learning to control the output of language models is certainly a form of AI literacy, the need for “prompt engineering” may be consigned in the history books to a curiosity associated with the earliest releases, as people sought to use the chatbot not just for conversation, but as a practical creative tool. A command line interface with highly unpredictable output is not the optimal user interface for co-creation.

How might the UX evolve ? Firstly, taking inspiration from the music revolution, I anticipate the emergence of writing environments will enable authors to orchestrate their writing in new ways. Perhaps the introductory user guide to an AI Writing Studio (Sept. 2023 Release) will describe functionality like this…

  • Source Apps. Select which AI writing generators you want to work with — the studio will render their drafts in different windows which you can arrange, refreshing them each time settings are changed. After a while you may figure out which ones work best with different styles, which ones are most responsive, or which ones are most fun!
  • Genre menu. Choose the genre of writing you want to work in (e.g., tech blog; journal article; business report; etc.)
  • Modulators. Configure the libraries of sliders down the side that you want — these remind you how the text can be modified and encourage experimentation, applied either to the whole document, or the selected text (e.g., length; formality; reader age; etc.)
  • Record On/Off. Turn on recording to log your studio session, enabling Replay and Analytics (see below).
  • Replay. Fast Forward/Rewind through your document’s timeline to revisit key moments (e.g., recovering the state of the modulators and each app’s draft at a given moment — you can branch your document and explore another version).
  • Analytics. AI is used to generate summaries of your usage of AI generators, such as how much you request, reject, adopt or adapt AI suggestions. This helps evidence your critical engagement with AI, which your course will have mentioned. Check if your assignment requires you to include the WAL (Writing Analytics Link) and/or the 1-page report.
  • Feedback Tips. The feedback panel uses the best research on writing to assist your writing skills. In addition to the Analytics, other tabs show you well established indicators of the clarity of writing, and the depth of your reflection and argumentation (varies with the genre you chose). The analytics are your springboard into our personally recommended Practice Exercises and Pro Tips…
  • Practice Exercises. These help you get the most out AI writers, while ensuring that you’re building your own writing and thinking skills.
  • Pro Tips. We curated some of the best videos from our elite writers, who walk through their writing practices with Writing Visual Studio.

At some point perhaps I’ll mock the interface up, and even get to build it. But these are just preliminary ideas — there are far more creative possibilities, introduced next.

Moving beyond AI as ghostwriter demands creative UX design

Glenn Kleiman helpfully discusses (with the aid of GPT) the roles that AI writers can play — as editor, co-author, ghostwriter, and muse. The panic around cheating focuses on AI as ghostwriter, and we will watch the inevitable arms race between AI generators and detectors play out. Policing is important, but not the only mindset we need to adopt. What might interaction with AI writers playing the other roles look like?

We find clues in the communities spanning both academia and the tech industry who’ve been working on Computational Creativity and most recently, Human-AI Co-Creation with Generative Models. Consider Ken Arnold and colleagues, prototyping generative AI that augments rather than automates human writers. The purpose of the AI is to prompt the author with questions, promoting more critical thinking and better writing:

A team at MIT and Harvard are exploring the addition of audio and visual cues for creative writers, along with text: this paper exemplifies the kind of detailed analysis of tool usage that the Writing Synth hypothesis requires:

Continuing with creativity, consider the extraordinary work of Sarah Schwettmann who shows in this keynote talk how generative AI for The Met renders images of cultural artifacts that fill the gaps between artifacts that have been discovered. Thus, using “Generist Maps“, we can ask what might have happened if two cultures had met?

If this is possible for images, then is it not plausible to envision AI writers drafting texts in the interstitial spaces between known evidence, or between polarised positions in an argument? And might visual interfaces such as the above not be an engaging way to work with drafts ?

In another project, Schwettmann’s team describe the Latent Compass, a prototype to help individuals generate images that match their intuitions about the meaning of complex terms (e.g., “more festive”, “more inviting”).  In a writing context, I can see authors teaching their AI assistant with examples of what they mean by “the crisis motivating a research program”, or “artful critique of an argument”, which become new, user-defined modulators. I would see such developments as evidence for the writing synths hypothesis.

The Human-AI co-creation community also offers us conceptual language to describe how the human and AI can be configured to co-create together, such as this example from Michael Muller and colleagues:

[Update 27 Mar 2023] Most recently they have proposed a set of user-centred design principles to guide generative AI tools, which I am looking forward to thinking through in relation to the writing studio concept:

So, those are glimpses of where we may be heading with human-AI co-writing. But right now, in the true Silicon Valley ethos of “move fast and break stuff”, we are witnessing the largest scale introduction of AI in education, with no evidence of its utility for learning. And it is in the trenches of everyday education, at school and university, and professional learning in the workplace, where the writing synth hypothesis must be tested.

Teaching and assessing writing: many hopes and fears, little evidence

GenAI is a system shock because teaching and assessment regimes rest on the assumption is that the learner has written the text, and that the goal is to assess their ability to do so unaided by anyone or anything else (other than the passive capabilities of word processors). Learners may, of course, draw on others’ work, but only following well-established guidelines (e.g., through quoting and citing), in order to maintain academic integrity. Some students cross the line into the territory of student misconduct (the reasons for which are complex), and an array of policies and software products to police this are in place.

However, rather than simply banning AI writing, this has also triggered an outbreak of creativity as educators share and debate ways to actively embrace the new possibilities of composing with AI, as a new way to cultivate students’ critical faculties. This is in my view absolutely the way forward. Our graduates must know how to orchestrate these instruments and (to borrow an aviation metaphor) fly them within their ‘flight envelope’ — understanding the limits within which they can be trusted to perform reliably, before the wings drop off…  Beyond that, if they’re to find work in the creative professions, students must be able to show the additional value that distinguishes them from 100% AI-generated writing, or mere AI app operators who can simply click buttons.

In my own university we are advising effective ethical engagement, and resourcing academics for shorter and longer term adaptation of their assignments, and most other universities are doing the same. This is a holding pattern while we wait for the dust to settle. The web is full of proposals for engaging students in using ChatGPT creatively (101 ideas), while others warn of the death of thinking. These myriad hopes, fears and advice are filling the vacuum of evidence at this transition point. In a year’s time we’ll have many anecdotes and practitioner reports, and the first robust peer reviewed research evidence.

However, while generative AI is undeniably new, we are not in completely uncharted waters. AIED research has been under way for over 40 years. There are communities dedicated to prototyping and evaluating computational support for writing, conversational user interfaces and pedagogical agents, to name just three at the intersection of ChatGPT as a design concept. The media conversation would be more informed if these researchers can translate their work into accessible forms for wider audiences, as well as apply their expertise to show how generative AI can be designed and deployed in ways that respect with what we already know. We’ve made a start on that conversation in my own institution.

[Update 14 Aug 2023: An expert forum on the future of assessment in the age of AI just wrapped up, and the report will be shared in this TEQSA webinar]

Knowledge, skills and dispositions for critical engagement with AI

Amidst all the excitement among the optimists, let’s consider one of the most prevalent aspirations: that students will critically engage with AI draft writing, identify its weaknesses, and show how they have improved on it in their submitted work. While academics proposing these ideas are able to do this, I wonder if they overestimate their students’ knowledge, skills and dispositions to do so.

  • Curriculum/domain knowledge is needed to validate factual claims and spot significant omissions.
  • Rhetorical analysis and writing skills are needed to improve on prose which may in fact exceed many students own ability
  • Dispositions such as the curiosity and authenticity are needed to resist the temptation to just run with what the AI served up.

These qualities must be demonstrated rather than assumed, and educators should design for wide variability among their students in their capacity to critically engage. This is just one example of the evidence that needs to be gathered. [Update 8 June 2023: after 1 semester teaching with ChatGPT, we have initial evidence]

[Update 27 Mar 2023] The Academic Integrity debate around generative AI is (understandably) skewed to the ‘dark side’, and badly in need of more sophisticated vocabulary to talk about what it means to write with integrity with AI. Katy Gero’s exciting research illuminates how creative writers feel about AI writing aids, and is exactly the kind of work we need now. I’ll highlight just one aspect of her work, around the differing ways that writers feel about “authenticity”:

“Writers talked about authenticity, or their ‘voice’, as a concern when it came to incorporating the ideas or suggestions of others. Here, we describe four types of authenticity issues that came up in our interviews: 1) the reader’s sense of authenticity, 2) the impact of viewing suggestions, 3) differing opinions on where authenticity lies, and 4) human v. computer authenticity issues.” (Gero, 2022: p.106)

Gero’s work may offer us concepts and language to help students develop their own sense of what it feels like to work authentically with AI writers.

Writing analytics and academic integrity

(An earlier version of this section was originally posted here)

Recall Analytics in my envisioned writing studio. In the near future, GenAI will be fully integrated into interactive tools for writing, coding, and other creative work with image, music, animation, video etc… I envisage our students will become power-users. Human-AI interaction ‘flow states’ will become a synergistic blur, as prompts are invoked by the learner or offered by the machine, and rejected, adopted, adapted — each in the space of a few seconds. Tens of thousands of times in the production of an assignment.

Asking a student to “declare/document what role AI played” after hours/days/weeks of working in close partnership with such tools now becomes an impossible question to answer.

Instead, following the Writing Synth hypothesis, we look to the music world and borrow a studio recording session analogy: we immerse ourselves in our work, it’s all being recorded, and then we need to replay and review, dissect and debate, re-record elements, or start over…

In educational terms, such tools will be scaffolding “reflection-on-action” (Donald Schön) by the learner, possibly also with peers, and the teaching/coaching team. In time, they develop the capacity to engage in increasingly nuanced “reflection-in-action”, making improvised decisions about how and when to call on AI…  In the language of human-computer interaction research, such tools will support Retrospective Cued Recall.

Analytics crunching that data will make visible patterns that are useful for improving performance. My colleague Antonette Shibani has already prototyped this (see below). We will be able to see — literally — how virtuoso performance with such tools differs from less developed performances. This can serve as formative feedback to the learner, and assist should academic integrity questions arise.

[Update 26 Apr 2023] Antonette Shibani, et al (2023). Visual representation of co-authorship with GPT-3: Studying human-machine interaction for effective writing. 16th International Conference on Educational Data Mining

The fundamental question, then, is whether students are learning to produce great work. And in the future, great work will not be merely what can be automated. As Michael Feldstein has noted, students must learn the limits of GenAI, so that they develop the qualities needed to produce work that is beyond full automation — and stay employed.

And so we return to assessment.

If you can’t write without AI, can you really write?

In a prescient paper written at the turn of the 90s, Gavriel Salomon, David Perkins and Tamar Globerson considered critical educational questions that they envisaged arising with “intelligent technologies” as they termed them. When we ask what effect AI has on students, they distinguish between performance with the AI, and the effects of using AI on the student, assessable once the AI is removed. Intriguingly, they invite us to imagine a positive, futuristic scenario:

“For another illustration, consider the possible impact of a truly intelligent word processor: On the one hand, students might write better while writing with it; on the other hand, writing with such an intelligent word processor might teach students principles about the craft of writing that they could apply widely when writing with only a simple word processor; this suggests effects of it.”

Well, here we are! Fast-forward 30 years to today, and some have argued that ChatGPT is an educational disaster because we only learn to think by writing (Rob Reich, p.20). Decades of research into writing does indeed show that the writing process activates many cognitive faculties for critical thinking. But the roles that a conversational, generative AI agent can play in provoking deeper thinking (see above examples) are not taken into account by such cognitive models, which assume a solo author.

Salomon et al. argue for mindful versus mindless engagement with AI to achieve high performance with AI, and pose the assessment question now confronting us today: should we evaluate what a student is capable of when using AI to augment their intellect, or the “cognitive residue” as they term it — how well they perform once stripped of the AI ? For many educators, it would be a dereliction of duty to turn out graduates who could not write well with a pen and paper, while for others, that is to be stuck in the past. The imperative is to graduate capable of high performance with a profession’s state of the art tools. It may of course be a false dichotomy if the latter is impossible without the former, but that is an empirical question.

Writing in the early 90s, pre-Web, pre-mobile, pre-Big Data, and pre-LLMs, the authors conclude that we cannot afford to assess only AI-augmented student performance. After all:

“Until intelligent technologies become as ubiquitous as pencil and paper—and we are not there yet by a long shot—how a person functions away from intelligent technologies must be considered. Moreover, even if computer technology became as ubiquitous as the pencil, students would still face an infinite number of problems to solve, new kinds of knowledge to mentally construct, and decisions to make, for which no intelligent technology would be available or accessible.”

We might question this assumption now — but something deep inside us as educators might whisper that we will have really lost the plot if our graduates cannot function without computational support. The resolution may lie in what exactly we want students to bring. Rose Luckin and Margaret Bearman have argued that it is pointless to assess students on anything that AI can do better, which is a rapidly rising waterline (Salomon et al. contest this). I’ve also argued that we need to move to higher ground and  harness analytics and AI to help where they can in cultivating the qualities and capabilities that are still distinctively human. How about we start with dignity, compassion and justice.

To close…

So, that’s the Writing Synth Hypothesis. I had fun writing it — let’s see how it all unfolds. This is indeed an extraordinary time.

Your comments are most welcome: my blog doesn’t have great discussion tools, so join the conversation in this LinkedIn thread.

Compendium Archive & Network 2021

CONTEXT… Those of you who know my R&D will know that I have a long-standing interest in the role that software can play in human sensemaking around wicked problems, a form of “augmenting human intellect” (Doug Engelbart). A powerful example is structured, visual hypermedia that makes tangible the ways that ideas, data and arguments connect with each other and documents.

This is certainly a story about the evolution of an interactive visual tool for thinking—but far more interestingly, it’s about the co-evolution of software with a set of practices to develop human fluency with the tool. Those practices were (in historical order) Dialogue Mapping, Issue & Argument Mapping, and Knowledge Art.

Here’s how I told this story in 2014, tracing my work back to Engelbart’s vision of personal computing and collective IQ, and here’s a quick overview of key books and papers.

COMPENDIUM… provides extremely flexible hypermedia linking between conceptual objects (e.g. questions, ideas, arguments), data and documents (local/online), through a visual user interface. As a hypertext system for thinking with, connections can be made in multiple ways: spatial proximity and visual linking within a view, plus tagging and transclusions across views (i.e. embedding a node in multiple views). Views can contain each other non-hierarchically (A can contain B which can contain A). Search can be refined by node types and tags. This 2014 RAE Impact Case distills the research –> impact narrative.

It is an open source, desktop Java application with a full SQL database, with XML and SQL import/export, and HTML publishing. In the course of its development, it was interoperable with (at the time state of the art) Jabber (open source instant messaging) XML, and semantic web RDF.

HYPERTEXT HISTORY… Compendium is descended from the pioneering hypermedia system gIBIS (graphical Issue-Based Information System) at MCC Labs led by Jeff Conklin and Michael Begeman, which led to the commercial corporate memory product CM1, renamed Questmap. Compendium was then developed over the course of around 20 years R&D starting in the early 90s at Bell Atlantic labs (White Plains NY) led by Al Selvin and Maarten Sierhuis, continued at NASA Ames Research Centre by Maarten on the Mobile Agents project providing a human-agent science team tool, in collaboration with my team at the Open University’s Knowledge Media Institute (KMI) from 1995-2014. The gIBIS/CM1/Questmap/Compendium lineage exemplifies an influential strand of hypertext R&D that preceded the invention of the Web, which like Xerox NoteCards, exemplifies what Frank Halasz called hypertext for idea processing — focusing on visualising and managing the connections between nodes as a form of intellectual work. You can learn a lot more about the intellectual lineage of these ideas in this brief history, another account, on this blog via the compendium tag, and on the memorial archive of Compendium co-inventor, my colleague, PhD student and friend, Al Selvin.

COMPENDIUM INSTITUTE… The Compendium Institute coordinated the international user/developer network, feature requests, code releases, research and training. Software development has been on pause since 2013, and may well not be continued, since much of the world now expects web-based tools (indeed see the great work by DebateGraph, and the KMI team developed quite a few). However, there remains an active user base of people who value the speed and functionality of a Java desktop app which continues to run on current Mac/Win/Linux Java.

I have therefore archived the Compendium Institute website for posterity, since it contains lots of resources:

SOFTWARE… You want the tool! I’m trying to keep links alive, and Java seems to be holding up remarkably since development ended around 2016.

  • If you’re a developer, then you can get the most recent CompendiumNG, from the CompendiumNG (next generation) wiki.
  • I can also offer this Mac installer version of CNG which requires no code knowledge to get running
  • If you’re non-technical on Win/Linux then this version of Compendium comes in an integrated installer:
    • KMi Open University version 2.0beta with advanced experimental features (like Maps supporting video annotation): Downloads page | QuickStart Guide for Mac or Windows *follow the guide*
    • CogNexus Institute also offers download links for a slightly older version

COMMUNITY… Since Yahoo closed down their groups end of last year, I’ve created a new Google Group which you’re warmly invited to join if you want to stay connected with fellow users and some of the original team. Collectively, we will hopefully be able to answer any queries about Compendium’s functionality and design rationale — and who knows, possible futures…

2020: strengthening the Quantitative Ethnography community

2020 will be remembered for many things… but amidst the disruption, it’s been a year of consolidation for the exciting, emerging field of Quantitative Ethnography (the book by David Williamson Shaffer; my review for Jnl. Learning Analytics).

The newly launched International Society for QE has been coordinating virtual events to strengthen professional ties across the globe, upskill researchers in the new tools and techniques, and the 2nd international conference is in Feb 2021. ISQE are to be congratulated on this progress, and in particular, the Epistemic Analytics Lab at U. Wisconsin-Madison are doing an awesome job in generously sharing their expertise, and making their work available through free analytical tools.

I was honoured to be asked to help design and chair the monthly webinar series which is building a library of examples how QE methods can be applied in diverse contexts. That’s proven to be a fascinating experience, and our own work (based on Vanessa Echeverria’s PhD) wrapped this up earlier this month (more coming in 2021!).

Abstract: Collocated, face-to-face teamwork remains a pervasive mode of working and learning, which is hard to replicate online. In team-based situations, learners’ embodied, multimodal interaction with each other and with digital and material resources has been studied by researchers, but due to its complexity, has remained opaque to automated analysis. The ready availability of sensors makes it increasingly affordable to instrument work spaces to automatically capture activity traces to study teamwork and groupwork. Yet, a key challenge is the enrichment of these multiple and intertwined quantitative data streams with the qualitative insights needed to make sense of them. In this seminar, we will discuss our inroads into giving meaning to multimodal group data. We have followed a human-centred approach to design meaningful end-user interfaces that convert multimodal data into data stories. Based on Quantitative Ethnography principles, we developed a modelling technique, termed the Multimodal Matrix, to grounding quantitative data in the semantics derived from a qualitative interpretation of the context from which it arises. We will present practical examples in the context of high-fidelity clinical simulations in which multimodal data (physiological, positioning, and logged actions) have been transformed into learning analytics interfaces that support teachers’ and learners’ reflection.

Video: Transcript

Papers:

The Multimodal Matrix as a Quantitative Ethnography Methodology. Advances in Quantitative Ethnography.

Towards Collaboration Translucence: Giving Meaning to Multimodal Group Data. Human Factors in Computing Systems.

From Data to Insights: A Layered Storytelling Approach for Multimodal Learning Analytics. Human Factors in Computing Systems.

Presentation: Slides

The craft + tech of structuring participatory deliberation

RSA kickoff webinar on deliberation

As a Fellow of the RSA I’m happy to draw attention to the important new series of webinars just launched, on the critical role that effective, participatory deliberation has to play in resolving complex challenges, even apparently intractable dilemmas — at many different scales, from an organisation, to a local community, city, regional or even international scale.

As happens sometimes, I ended discovering a colleague at my own university doing fantastic work! Check out Nivek Thompson and her Deliberately Engaging portal. In prepping some notes for her, I thought I might as well blog them in case of wider interest to this community.


Hypermedia Discourse

A lot of my work has investigated a particular way in which software can help make thinking visible, the focus of all my work. Such tools seek to “augment human intellect” in Doug Engelbart‘s memorable words (my tribute to his inspiration for my work, and what he thought about this [Visualizing Argumentation]).

Here are some examples of how this works:

  • Make aspects of the conversational structure visible. Once a phenomenon is visible, rendered in a visual language that provides helpful ways to reflect on what is unfolding, it can be talked about, and is an “improvable object”. The Hypermedia Discourse project prototyped and evaluated the potential of combining models of dialogue and argumentation, with hypertext functionality for connecting issues, ideas, arguments and documents. Since we were interested in discourse about wicked problems, differences in perspective were the default starting point.
  • Support online forum moderators/facilitators assess the health of the conversation. A well designed user interface helps online participants to structure their contributions in ways that can provide the software with new ways to check the state of the debate (not possible with conventional flat chats, or threaded forums), and reflect this back to participants and/or moderators (see the Catalyst project for example).
  • Help to track ideas. Hypertext systems (more powerful than the Web) provide flexible ways to keep track of ideas (nodes), not just information. Anna De Liddo’s doctoral research is an example of how this can provide new forms of accountability in the participatory process ((in her work, for participatory urban design).

An important strand of our work was examining the facilitator skillset and disposition required to make good use of visualizations in real time, to augment the deliberation. Spearheaded by Al Selvin’s doctoral research, this led to a book that set out the concept of Knowledge Art.


Collaborative Evidence-based Problem-Solving

More recent work led by Tim van Gelder at Melbourne University (an Argument Mapping philosopher and software entrepreneur) has broken new ground in a particular niche of the design space: how do you convene a team of citizens to tackle a complex problem, with the challenge of devising an evidence-based, plausible analysis of the best way forward?

They have just published exciting results demonstrating that some teams of volunteers recruited via Facebook performed as well as, and in some cases better than, teams of professional intelligence analysts. See the paper to appear in the Journal of Cognitive Engineering and Decision Making on the Hunt Lab website, and this CIC webinar.


The emergence of NLP to detect critical, reflective writing

Natural Language Processing (NLP) has in recent years emerged from the AI labs into the mainstream. This has been a recent focus of my work, in the context of giving students instant feedback on their drafts. This has yet to be deployed in the context of participatory deliberation, but here are some preliminary reflections on where the automated detection of shallow and deeper reflection might assist participants posting online to reflect on how they are reacting to challenges — from other people, or the turbulent life events that are threatening dearly held assumptions, and ways of life.

Might the growing potential of NLP to make sense of rich, narrative prose offer the optimal combination in years to come — playing to the respective strengths of machines and humans to make sense of the world?


Power tools (and new literacy?) for deliberation professionals?

I remain excited about the potential of interactive, usable visualizations to help tackle the limitations of individual and collective human cognition. I have also seen first hand how hard it is for people to learn to structure their thinking more carefully than firing off their thoughts in the usual way. The role of the deliberation facilitator can be absolutely critical to modelling and scaffolding stakeholders into more reflective modes of reflective dialogue and rigorous argumentation.

That provides the basis for a good conversation with the growing international networks of deliberation experts who will also be working increasingly in online or hybrid modes.

Should predictive models of student outcome be “colour-blind”?

This post was sparked by the international condemnation of George Floyd’s death, and the many others who came before him. Many communities and institutions are now reflecting on how structural racism manifests in their work (e.g. see SoLAR’s BLM statement and resources to help members learn more).

This is a tentative step into issues of race, about which I should declare I have no academic grounding. Nonetheless, it is important to ask what the implications are for a specific form of Learning Analytics, namely the predictive modelling of student outcomes. Should demographic attributes such as ethnicity be explicitly modelled, or should the models be “colour-blind”? While all categories have politics, this struck me as an interesting question, given that such techniques are demonstrating their value specifically in levelling the university playing field for all students. 

With thanks to Madi Whitman, Bart Rienties, Marti Hlosta and Paul Prinsloo for initial fact-checking and feedback. All comments are welcomed via this blog (moderated), the twitter thread or the LA Google Group thread.


Be more white. Be more male. Be wealthier. Those are the biggest correlations with success. It’s terrible, but it’s the truth.
[12] (p.1)

Classification systems provide both a warrant and a tool for forgetting […] what to forget and how to forget it […] The argument comes down to asking not only what gets coded in but what gets coded out of a given scheme.
[13] (pp. 277, 278, 281)

Since the emergence of Learning Analytics (c.2011) as both an intellectual community and commercial marketplace, an influential strand of work in higher education has been the use of predictive analytics, that is, developing computational models to identify students who look statistically likely (i.e. on the evidence of similar past cohorts) to be struggling, at risk of failing, or even dropping out. This is a dominant form of analytics inherited from the business world and machine learning, where it is highly lucrative to be able to predict the likelihood of, for instance, a customer buying a product or switching service provider — and take anticipatory action to change that possible future. So why not do the same for education?

Debate surrounds the ethics of such models in higher education, a particular version of broader concerns around the “datafication” of education through analytics, and now AI. The issues are complex, but examples of constructive dialogue are emerging, in which Learning Analytics and AI in Education engage with such critiques (e.g. these recent edited collections [2-4]).

Predictive modelling intersects with questions around the profiling of students, one attribute being ethnicity, which is what I want to focus on here given the current times we’re in, just a few weeks after the death of George Floyd at the hands of the police.

High profile success stories serve as iconic posters for the use of predictive modelling of student outcomes. Consider the Georgia State University Graduate Progression Success Advising program. It’s not called GPS by accident: the predictive model alerts student support teams when students look like they’ve ‘missed a turning’ (to push the metaphor) and off-course. An example screen from the system is shown below.

Discipline-level, cohort summaries of Low, Medium and High risk levels in the Georgia State University Graduate Progression Success Advising program.

Intriguingly, with regard to the question of racial colour-blindness, there’s a strong social justice angle that challenges head-on the demographically-related achievement gaps that many universities know only too well. Tim Renick, VP (Enrollment) at Georgia State University is unapologetic about GSU’s mission, and the GPS Advise website proclaims the sophistication of the analytics that help to power this:

“We have eliminated achievement gaps. For the last four years, we have been the only national university at which black, Hispanic, first-generation and low-income students graduated at rates at or above the rate of the student body overall. Georgia State is showing, contrary to what experts have said for decades, that demographics are not destiny.

Students from all backgrounds can succeed at comparable rates. Predictive analytics have helped all demographic groups graduate at higher rates from Georgia State, but just as critically, they have helped to level the playing field for all of our students.”

The irony will not be lost on those concerned about the datafication of education. Here we have analytics helping to level what historically has not been a level playing field for all students. When tools such as this are used intelligently, as aids for student support teams who are very much in the intervention loop, producing impressive outcomes for historically minoritized groups such as these (evidence which is not contested to my knowledge) — well, what’s not to like?

Another mature example of the process of embedding a predictive modelling tool into work practices is from The Open University UK (webinar / paper / paper [6, 7]). Working with online distance learning students, most of them mature students returning to academic study long after leaving high school, and including a high proportion of students with accessibility needs, the OU team has shown that compared to staff who did not use OU Analyse to monitor student progress, those who did contacted them more, with higher success rates [5]. Again, here we have analytics helping traditionally disenfranchised cohorts.

A screenshot from the OU Analyse dashboard, showing the risk of each student not submitting an assignment, their predicted grade, and their probability of passing or failing the course. (Figure 2 from [7])

Having set the scene, I want to focus on a specific decision that has to be made in such work, which I’m framing as follows:

Should predictive models of student outcome be “colour-blind”?

Two sides of the debate go something like this:

YES: MODELS SHOULD IGNORE HISTORIC INJUSTICES. Predictive models should ignore demographic attributes, which are well known to be highly predictive of outcomes, but students obviously have no control over their ethnicity, high school, being first-generation-in-family at university, etc. It’s clearly unethical to classify students as higher risk from day 1 for those reasons, immediately placing them in the shadow of inequitable historical patterns. They’ve got to university, possibly demonstrating greater resilience than their more privileged peers, so we wipe the slate clean. What counts is what they do when they walk through the door, some of which can be tracked by analytics through digital activity traces. Such models can therefore be declared to be “colour-blind”: ethnicity is not modelled explicitly, and nor are any other known proxies (e.g. Zip code; High School).

NO: MODELS SHOULD REFLECT BUT NOT PERPETUATE ALL KNOWN FACTORS. Predictive models of student success/risk should include demographic variables, since they greatly improve the model’s performance. It is myopic to ignore this, just as we should not ignore science and social science when they provide solid evidence of other difficult truths about societal inequities. The student’s demographics are not held against them, but rather, used to improve their chances. We should thus model student risk as comprehensively as possible, with our ethical ‘eyes wide open’, forearmed to use this knowledge in the students’ best interests, with strong ethical principles to ensure that competing interests are not allowed to influence decisions (e.g. a student’s need for extra support has resource implications).

Until recently, I thought of these positions as rather polarised. But a third analysis struggles with an unequivocal yes or no. This view problematises the goal of even trying to achieve colour-blindness:

BEING “COLOUR-BLIND” ≠ BEING ETHICAL

I’ll state very clearly that I’m brand new to reading anything academic about racism. As a result of reading sparked by George Floyd’s murder, I only just became aware of the work of people like Eduardo Bonilla-Silva on the nature of white privilege and structural racism, and at this point, have only managed to read various summaries and reviews of his influential book, Racism without racists: Color-blind racism and the persistence of racial inequality in the United States [1]. He argues:

“Whereas Jim Crow racism explained blacks’ social standing as the result of their biological and moral inferiority, color-blind racism avoids such facile arguments. Instead, whites rationalize minorities’ contemporary status as the product of market dynamics, naturally occurring phenomena, and blacks’ imputed cultural limitations” (p.2).

“Much as Jim Crow racism served as the glue for defending a brutal and overt system of racial oppression in the pre-Civil Rights era, color-blind racism serves today as the ideological armor for a covert and institutionalized system in the post-Civil Rights era” (p.3)

Colour-blind racism operates through:

  1. liberalism (markets are open to all and do not discriminate)
  2. naturalization (people “naturally” segregate themselves from other racial groups)
  3. cultural racism (minorities participate in self-defeating behavior) and
  4. minimization of racism (racism is no longer prevalent to address, specifically).

I found another article fascinating, introducing critical race theory to reflect on how academia functions, specifically HCI, a sister field to Learning Analytics (which just won CHI’20 Best Paper) [9]. In their summary of critical race theory, the authors also note Bonilla-Silva’s point (1) above:

“Liberalism itself can hinder anti-racist progress [34]. Liberalism’s very aspirations to color-blindness and equality – while admirable – can impede its goals, as they prohibit race-conscious attempts to right historical wrongs. In addition, liberalism’s tendency to focus on high-minded abstractions can lead to neglect of discrimination in practice.” (p.3)

These ideas raised the question in my mind: does making our computational infrastructure “colour-blind” merely perpetuate systemic discrimination in universities? So I was delighted to read the work of Madi Whitman [12], who presents an ethnographic account of how a university made its modelling decisions. There are some interesting quotes from the data science team, which I suspect might be echoed by many others, who are trying to make ethical decisions. First they are aware of the uncomfortable truth, as are many universities:

“Be more white. Be more male. Be wealthier. Those are the biggest correlations with success. It’s terrible, but it’s the truth.”

—Excerpt from interview with Don, a university administrator [12] (p.1)

Since the predictive model drives automated nudges to the students, they try to do the right thing (for the YES camp) — exclude demographic attributes over which students have no control:

“Socioeconomic status things. Demographic markers. But they’re all things that either because it’s too late in the game, we can’t tell a student, “Boy, it would have been great if you would have studied harder in high school.” And we certainly can’t tell a student on a demographic or socioeconomic thing, we can’t say, “Hey, it’d be good if you weren’t so poor.” There’s nothing a student can do with that. Even though it does put ‘em in a higher risk category. So we took those things that were malleable by the students. Things like, how much time they were spending on campus. Whether they were a proxy for whether we believed they were paying attention in class by how much data they were downloading in a class.” (p. 6)

Note the strong argument for student agency, which is a principle valued in much ethical discourse in Learning Analytics, and Human-Centred Design thinking. The student should be in control:

“I guess that we assume that what [students] did in the course of the day, they had control over. Right, so they chose whether they were gonna eat or not . . . they chose the gym or not, being on campus or not . . . They chose living where they chose to live. I think they have some say in that…So it seemed to me that any time that they had an opportunity to make a decision about what they were going to be doing, we called that a behavior.” (p.7)

Whitman helps us understand that while the analytics team sees this as the ethical response, it’s a double-edged sword: do they really have that level of control? She argues that:

“Because attributes are removed from the model and nudging, the reliance on behaviors suggests that students’ choices are at the heart of their success at the institution. Because demographic data are not incorporated into the predictive model at all, success is linked with behaviors and students’ choices. The purposeful presentation of data to students encourages students to internalize those data and act on them. As such, responsibility now rests on the students to take hold of their success.” (p.10)

If you are in the YES camp, this is exactly the goal. Level the playing field, we don’t care what colour you are, everyone is must take responsibility for their study habits, level of engagement, assignment submission, etc.

However, might this not also resonate with items 1, 3 and 4 in Bonilla-Silva’s work introduced above? The university and its learning platforms are framed as “open markets”, with opportunity for all (1); if students do not make wise choices, they only have themselves to blame (3), because we’ve erased racism from the algorithms (4):

  1. liberalism (markets are open to all and do not discriminate)
  2. naturalization (people “naturally” segregate themselves from other racial groups)
  3. cultural racism (minorities participate in self-defeating behavior) and
  4. minimization of racism (racism is no longer prevalent to address, specifically).

So Whitman with her modelling case study, and Bonilla-Silva in general, are questioning whether students from historically marginalised groups are really as autonomous and agentic as their more privileged peers. Whitman concludes:

“The visualizations of certain kinds of data—namely data students ought to use to inform their everyday decision-making—and obscuring of demographic data place the burden of responsibility and success on students. By minimizing the role that race, class, and gender play on graduation outcomes, the institution, through the model, can present behaviors as major factors in the likelihood of a student grad- uating within four years. If students do not attend class, a low GPA is a consequence of that decision.

Thus, the constraints around choices become invisible. The university and its existing inequalities start to vanish because success is placed in the hands of students. Social climate problems, structural barriers, issues of belongingness, and resource shortages disappear. A student cannot cite external factors in this model of success dominated by behaviors. The result is a shift in a locus of responsibility, wherein nudging is meant to give students tools to manage themselves and regulate their own behavior based on insights they ought to draw from their data.” (p.10)

WAYS FORWARD?

There seem to be some questions that could be asked, as a way to move this forward.

Does anyone contest the positive outcomes for students from the use of predictive models?

For instance, when GSU reports the startling impact of the GPS Advising initiative, is anybody questioning the figures? Is anyone questioning the claim that the algorithm has a pivotal role to play in this, rather than the impressive level of human support available to students? At the Open University, we knew that simply calling a student increased the chances of a positive outcome.

What is the purpose of the modelling?

If you’re designing automated nudges for students (as in the Whitman case study), clearly, there’s no point nudging them based on their static demographic history, so removing such attributes from the model seems uncontroversial in modelling terms. Whitman, of course, is concerned about this erasure (but see next section as to whether this is justified).

If you’re designing a model to understand the spectrum of challenges students face, in order to understand how to support them, then ignoring demographics becomes problematic. The UK Open University’s Student Probability Model  [7] was developed for financial forecasting, assessing the likelihood of a student still being enrolled as the course unfolded (sometimes over years for part-time students). This took into account deprivation indices, which could of course be a proxy for race in some contexts, but erasing this would simply lead to more erroneous financial forecasts. We should ask (perhaps even more so in these straightened times for universities) if it is in anybody’s interests for universities not to budget as accurately as possible.

The OU Analyse predictive model also takes into consideration a range of demographic variables including socio-economic and ethnic when making the first initial predictions, before a course starts. However, nearly all of the demographic factors quickly lose relevance once actual engagement and behavioural data is gathered when a course begins, in particular once the first assessment deadline has passed. Furthermore, previous credits obtained is mostly more predictive than any demographics. Interestingly, while the OU Analyse team has wanted to remove demographics given the limited additional variance its explains, those teaching on the front line apparently prefer to retain this, since it helps them to ‘colour in’ their picture of a student. Ethical arguments for both the Yes and No camps?

Given this tension between quant and qual drivers, it seems particularly important to understand when and why predictive models fail, through close qualitative analysis (see this recent example from the OU team [8]), as well as to understand in detail the experiences of the student support teams who use – or are expected to use – the outputs predictive models (e.g. [5]).

Is any real harm is caused by colour-blind modelling?

Whitman argues that in principle, an unfair burden is imposed on marginalised students if we assume they have the same capacity as their more privileged peers to respond to nudges and make wise choices. There is plenty of evidence that marginalised groups are not as free to make the same life-choices as more privileged whites, but is there any empirical evidence yet regarding student choices in response to automated nudges? I don’t know any yet.

One size does not fit all: students with the same demographics may still be very diverse

A black student may be working from home, in very poor physical and emotional conditions, poor computing and network access, struggling financially, commuting long hours, with dependents to care for. That student is clearly battling constraints that others are not, which will seriously affect how much “control” they have over their choices, through no fault of their own.

  • This is all invisible in the colour-blind model (YES camp). It is visible when we model such metadata (NO camp) and could be taken into account.

Another black student may have a generous scholarship, living on campus, free from carer responsibilities, and able to seize every opportunity that comes their way.

  • This seems to be the default assumption behind colour-blind student modelling — and that is precisely the point.

Should we just stop using predictive models in education?

Despite the flagship examples, perhaps the potential for poorly implemented predictive modelling is so high that they’re best steered clear of. It’s complex both technically and ethically. A range of ethical concerns not covered includes:

  • One size does not fit all. A body of evidence now demonstrates that a predictive model for one course does not translate smoothly to other courses. Differences in discipline, cohort, pedagogy and learning design introduce myriad variables.
    But within a given course, things are simpler, surely?
  • We don’t necessarily want to teach the way we always have. Predictive models assume that historically stable patterns are a reliable predictor of the future. But even within a course, this is not always true, since teaching staff, curriculum and pedagogies change. Indeed, many universities are trying to shift the way their staff teach and assess to more future-focused pedagogies. Innovations by definition break from the past, and so will likely break the predictive model, and the last thing we want is for our analytics to act as a brake on improving teaching. In our pandemic-afflicted world, predictive models based on a blended pedagogy with on-campus students, are unlikely to translate smoothly to 100% online students, working from diverse timezones (but that is ultimately, an empirically testable question).
  • Risk of misclassification. As in all areas of society where algorithms are classifying people, there is growing concern over the risk of being misclassified. Who wants a High Risk of Failure flag on their record, even before they start their studies? Is that flag really deleted, or saved to help validate future models? And could that classification be leaked to other entities, who could use it inappropriately?
  • University lacks the capacity to act. Prinsloo and Slade argue that a university has at least a moral, if not legal, obligation to act if it believes a student is at risk of failure. Predictive models, when valid, thus place a new burden on universities [11]. A key take-home from mature case studies such as GSU and the OU clarify the investment in people, processes and tools required to deliver on this.

So, there are significant risks that universities could buy predictive modelling products like any other ed-tech, but either use them badly, or if they are tuned well, still cannot act on what the dashboards are telling them, thus opening themselves up to charges of negligence. Perhaps it’s better not to know tens of thousands of students’ risk profiles in such precise terms…

Many universities choose instead to focus on other forms of analytics that make visible student activity in helpful ways, to both educators and students, provide educators with tools to intervene with personalised feedback at scale [10], but make no attempt to build a risk profile. That profile is left implicit, inferred by (hopefully well trained) student support mentors and educators.

What do students think?

I’ll close with this obvious question, but not one with any empirical evidence I know of. Let’s bring diverse students into the conversation and consult with them on these matters. Learning Analytics is beginning to introduce human-centred design methods that give a voice to students, and as with any co-design process, this requires learning, and listening, by all stakeholders. However, I do not know of any that engages students around predictive models in particular, and issues of race specifically.

How do students from diverse backgrounds engage with questions such as these?…

  • Do you want to be treated by the university just like any other student? Or should the university be recognising that you come from very different backgrounds, live in very different conditions, facing very different challenges day-to-day?
  • This extends into our IT systems: what do you think about analytics that continuously predict your likelihood of success, to maximise the support we can give you? Demographics including ethnicity and postcode can help improve such models, and help us ensure that outcomes are equitable for all students – does that seem reasonable? 
  • Are you surprised or shocked, or would you expect no less from a technically advanced university?
  • Are you happy to trust that the university will behave ethically, or do you want more transparency? How much do you want to know about the data we have and how we use it, and how much control do you want over this data?

References

[1] Bonilla-Silva, E. Racism without racists: Color-blind racism and the persistence of racial inequality in the United States. Rowman & Littlefield Publishers, 2006.

[2] Buckingham Shum, S. Critical Data Studies, Abstraction & Learning Analytics: Editorial to Selwyn’s LAK keynote and invited commentaries. Journal of Learning Analytics, 6, 3 (2019), 5-10 https://doi.org/10.18608/jla.2019.63.2

[3] Buckingham Shum, S., Ferguson, R. and Martinez-Maldonado, R. Human-Centred Learning Analytics. Journal of Learning Analytics, 6(2), 1–9. . Journal of Learning Analytics, 6, 2 (2019), 1-9 https://doi.org/10.18608/jla.2019.62.1

[4] Buckingham Shum, S. and Luckin, R. Learning analytics and AI: Politics, pedagogy and practices. British Journal of Educational Technology, 50, 6 (2019), 2785-2793 https://doi.org/10.1111/bjet.12880

[5] Herodotou, C., Rienties, B., Boroowa, A. and Zdrahal, Z. A large‑scale implementation of predictive learning analytics in higher education: the teachers’ role and perspective. Educational Technology Research Devevelopment, 67 (2019), 1273–1306 https://doi.org/10.1007/s11423-019-09685-0

[6] Herodotou, C., Rienties, B., Hlosta, M., Boroowa, A., Mangafa, C. and Zdrahal, Z. The scalable implementation of predictive learning analytics at a distance learning university: Insights from a longitudinal case study. The Internet and Higher Education, 45 (2020), 100725 https://doi.org/10.1016/j.iheduc.2020.100725

[7] Herodotou, C., Rienties, B., Verdin, B. and Boroowa, A. Predictive Learning Analytics ’At Scale’: Guidelines to Successful Implementation in Higher Education. Journal of Learning Analytics, 6, 1 (2019), 85-95 https://doi.org/10.18608/jla.2019.61.5

[8] Hlosta, M., Papathoma, T. and Herodotou, C. (2020). Explaining Errors in Predictions of At-Risk Students in Distance Learning Education. Proc. International Conference on Artificial Intelligence in Education (AIED 2020), pp 119-123. https://link.springer.com/chapter/10.1007/978-3-030-52240-7_22

[9] Ogbonnaya-Ogburu, I. F., Smith, A. D. R., To, A. and Toyama, K. Critical Race Theory for HCI. In Proceedings of the Proceedings of the 2020 CHI Conference on Human Factors in Computing Systems(Honolulu, HI, USA, 2020). Association for Computing Machinery. https://doi.org/10.1145/3313831.3376392

[10] Pardo, A., Bartimote, K., Buckingham Shum, S., Dawson, S., Gao, J., Gašević, D., Leichtweis, S., Liu, D., Martínez-Maldonado, R., Mirriahi, N., Moskal, A. C. M., Schulte, J., Siemens, G. and Vigentini, L. OnTask: Delivering Data-Informed, Personalized Learning Support Actions. Journal of Learning Analytics, 5, 3 (2018), 235-249 https://doi.org/10.18608/jla.2018.53.15

[11] Prinsloo, P. and Slade, S. An elephant in the learning analytics room: the obligation to act. In Proceedings of the Proceedings of the Seventh International Learning Analytics & Knowledge Conference(Vancouver, British Columbia, Canada, 2017). Association for Computing Machinery. https://doi.org/10.1145/3027385.3027406

[12] Whitman, M. “We called that a behavior”: The making of institutional data. Big Data & Society, 7, 1 (2020), 1-13 https://doi.org/10.1177/2053951720932200

[13] Bowker, G. C. and Star, L. S. (1999). Sorting Things Out: Classification and Its Consequences. MIT Press, Cambridge, MA.