Zencastr
00:00:00
00:00:01
Speed1x
Format
Share
Embed
Report

#35 Jiayin Zhi: When AI Helps Thinking—and When It Replaces It

AITEC Philosophy Podcast
AITEC Philosophy Podcast

27 plays · Aug 28, 2026

AI can speed up your work flow. But does it make you any smarter? In this episode of the AITEC Philosophy Podcast, Roberto and Sam talk with Jiayin Zhi, a PhD student in computer science at the University of Chicago, about what large language models are doing to our thinking. Her research asks a question that is becoming harder to avoid: when we use AI to interpret, write, and reason, are we extending our minds—or outsourcing them? The conversation begins with human-centered computing, the idea that technology should be designed around human needs, values, and real cognitive habits. From there, Jiayin shares findings from her research on AI-assisted close reading. One surprising result: AI-generated interpretations can improve the final written product, but too much AI can crowd out the pleasure of discovery and leave people feeling like there is little room for their own interpretation. The episode then turns to critical thinking. Here, timing matters. Using AI from the start—especially under time pressure—can produce polished work without deep understanding. But using AI later, after doing independent thinking first, may help reduce myside bias by introducing counterarguments and alternative perspectives. Along the way, we discuss poetry, interpretation, cognitive offloading, copy-and-paste learning, time pressure, myside bias, and the difference between using AI as a shortcut and using it as a genuine thinking partner. This episode is for anyone who wants to use AI well without surrendering the struggle that makes learning real. For more info, go to ethicscircle.org [https://www.ethicscircle.org/]. 

Transcript

Speaker: everyone and welcome back to the A-Tech Philosophy Podcast. Today we are joined by Jai-Yen Zi. Jai-Yen is a PhD student studying computer science at the University of Chicago.

Speaker: Her research is at the intersection of the fields of human-centered computing and artificial intelligence. Today we will be in conversation about her research and hopefully be getting actionable advice for how to profitably use large language models.

Speaker: Welcome to the A-Tech Philosophy Podcast. Yes, it's my pleasure of being here. ah Thank you for having me on this podcast. Sure, yeah. ah So the first two things we want to do ah right away is we want to get to know you. We want to know who you are.

Speaker: And I also want you to tell us about human-centered computing because I've been around the block and this is actually the first time that I hear that particular string of words put together, human-centered computing.

Speaker: So maybe while you're telling us about yourself, you can tell us what human-centered computing is and how you got into it. Yes, sure. So basically, human-centered computing is the idea that technology should be designed around people. like Basically, they're designed around people's needs, values, its so especially how they actually feel and think.

Speaker: ah So here, based on the design principle of human-centered computing, ah like it's like we want to build a technology system on that is more accessible for people to adapt to it. ah So here we start from like what do people actually need and feel and what is this doing to and what this technology is doing to them.

Speaker: And then we design technology system back from there. So broadly speaking, for my research, ah so I was initially motivated by the idea that I would like to design technology that can benefit people.

Speaker: But like on the other hand, of frankly speaking, for me, it actually started from my first-hand experience. For example, when I first learned about ai like when I first learned about these large language models, like ah in classes, especially about like on how they work under the hood and then like on how can we build a model from scratch.

Speaker: It's like, it feels like one thing, but like on the other of hand, when I started using these AI products myself, especially when I see those especially when I see those AI-generated content in the wild, and it actually felt completely different. So I couldn't help thinking about a question around like, what is this AI doing to like my own thought and then doing to my own thinking ability? And ah And so I think that it's not just for me.

Speaker: um So I think on one hand, now that we have built AI on with those great capability and and dealing like such technology has provoked some real anxiety about this impact on impact on human cognition. So on one hand, like AI is definitely changing how we think, but on the other hand, but we don't have if a clear evidence-based account of what AI is actually doing to our thinking. And then like, what can we do about such impact?

Speaker: on So what me and my research devoted to and doing is to building on distributing that evidence based account about AI's impact or thinking.

Speaker: ah Yeah, that's clear. like I've read a couple of your articles and in one of them it seems like you were thinking about the impact of AI usage on critical thinking and then another one you were thinking about the impact of AI usage on like interpreting cultural artifacts, something like poetry.

Speaker: Those are kind of like two areas where you've um research the impact, I guess, of AI on us. How would you describe overall what kind of results you found? You found that the negative, positive impact, what would you say there?

Speaker: and So i think like based on the series of experiments that I have conducted, I think like overall we cannot define AI's impact on just by one statement because it really depends on how the AI is actually designed and used. And then like it really it's like such AI's impact is really different based on the condition, which we can differ into.

Speaker: interesting yeah Yeah, one thing I can say right away is as I'm becoming familiar with your work, it's sort of obvious that, um I mean, it's becoming more obvious to me as I read your work that, you know, there's like two positions at the extremes, right? Some people don't want to use AI for anything because they think it's dangerous or whatever. Other people ah want to use it for everything. And so I think those two positions are obviously not helpful.

Speaker: And then you're left with the middle ground. And what your work is kind of pointing out is that that middle ground is extremely vast and complex. And depending on things like how much time you have ah ah or or you know how a task how difficult the task is, it changes how it is that we use the LLM and therefore how it is that what the LLM's impact on us is. Did I kind of capture your view?

Speaker: Yes, exactly. all right. Well, in that case, let's move into your ah first study on ah reading poetry and interpreting cultural artifacts.

Speaker: So in order to understand this, we need to have the listeners understand what close reading is. So ah I guess let's begin there. Tell us what close reading is.

Speaker: Yes, sure. Close reading is a skill that that most people learn ah started from their ah high school experience. So it is a skill that requires bit attention and then like it is the ability to understand, ah explain, interpret, and critique.

Speaker: cultural work. It's not just about what it says, but like why and how it says it. um For example, for close reading, you need to pay attention to things like word choices and then like the sound, ah like the form and then like the rhythm of the text and then ask how these choices were actually made.

Speaker: Okay. So It seems also that on top of that, you noted when you were explaining close reading in your article that you have to, it has to feel good. There has to be some pleasure associated with it. Now, Sam and I have degrees in philosophy and we can say that interpreting is not always pleasurable, but ah tell us why you think or ah why this construct of close reading involves the feeling of pleasure.

Speaker: Yeah, this is like this is ah interesting for me to know that like people can have, odd you guys may have a different experience about interpretation and close reading.

Speaker: From my perspective, it appears that um people's enjoyment and pleasure can play a pretty important role in close reading, especially on this feeling of pleasure especially arise from like this effortful engaging text discovery process of close reading.

Speaker: um For example, if we think about the process of close reading, like you may notice something first and then like you work out on like what the words like means yourself. And then like there might be a real reward in that process because of the textual insight that you just learned.

Speaker: And then there might also be the enjoyment of what you just discovered, like, for example, a line that surprised you. ah So I think it's also partially due to the satisfaction of on text discovery. And then and then like on the other hand, it's also about like being touched by the work itself, if that makes sense.

Speaker: I mean, that does, and I I do think typically, i don't know, we don't really necessarily need to go on a tangent on this, but I do think that um there's something to that point that i think most cases where you kind of unlock the meaning of a text or you sort of gain an insight into what a text is saying, especially with poetry,

Speaker: I do think it's generally pleasurable. When I think about counterexamples, they end up being more like stuff like, You can imagine someone who's maybe just like you know a long-time skilled poetry analyst, you know a literary critic who like could just do this stuff in his sleep.

Speaker: And at a certain point, you could imagine that he's able to just like figure out, you know analyze a text, do interpretive work in an almost like cold, mechanical way where he doesn't get much pleasure from it.

Speaker: But I, yeah, but I still think that like generally speaking, like honestly, when I think about my own experience, like most of the time, yeah, when I, especially with poetry, the thing is it's tricky with poetry is that there's generally objective value in the work itself. So that's like a varying condition, right? Like, I mean, if you're like, yeah, reading the poetry of like Nabokov, I mean, it's like great stuff. And so that could be relevant to why you're getting so much pleasure out of unlocking it whereas like what if you gave me like a really lame like legal sorry but um anyway it's like something like really boring and dry like a legal case like i don't know whatever whatever it is that people who are getting a jd study and like i could probably unlock things but would it be as pleasurable i mean maybe but

Speaker: Anyway, I'm sorry. I feel like I just gave some bunch of random yeah thoughts about it. but Maybe you can tell us, so if it isn't just poetry, so examples beyond poetry that might count as close reading. I do have but one question attached to that. I recently read Darwin's you know Voyage of the Beagle, and I was like enthralled, and I wanted to go back to South America and like revisit those.

Speaker: Does reading travelogues count as close reading? Oh, ah I think these are all great points. i think I can totally agree with that feeling. So I think like people's close reading experience ah can be really dependent on like ah can be really dependent on like what is the ah artifact that we read and then also like what's the goal of this reading. and then like it will really differ a lot ah like if we do close reading, say, for work or if we just do close reading for ah entertainment.

Speaker: And then like it will also differ a lot like based on the form of text. And then some examples beyond poetry for close reading is, like ah for example, like a catalog, as you mentioned, ah or like a story. And then like I think like a close reading as a social skill, it can also happen in other types of modality.

Speaker: For example, say a painting or like a speech or like a movie. So i think cloud based um I think based on my understanding, it it's like any rich artifact um that was asking the question about like how is this made and then what is this doing to me ah can be considered as examples oh for close reading.

Speaker: Yeah, well, I mean, I guess, yeah, we maybe we could dive into like the impact of AI on close reading. i mean it seems like your study of AI on close reading was really more evaluating the output.

Speaker: So it's like, i mean, you can kind of walk us through what exactly the person did, but my impression was like the person read um a poem, then some people kind of just generated an interpretation of the poem on their own. Other people sort of like,

Speaker: looked at an AI generated interpretation and then kind of like maybe used it to some degree to create their final interpretation. Then other people looked at like three different AI generated interpretations of the poem.

Speaker: But anyway, like it it seems like if you're just looking at the output, it seems like you're one of your findings was that, yeah, like people who kind of used generated

Speaker: interpretations of poetry, textual interpretations of poetry, which were better than people who didn't use any AI. And so, which I don't know, it could be because of the fact that people in our time don't read any poetry, but, but anyway, maybe you control for that. I don't, I'm not sure. So anyway, but anyway, yeah. So can you just kind of walk us through the results there?

Speaker: Yeah, sure. so So basically in this paper, ah we ran like this ah randomized experiment with about 400 participants because um ah And it's like, in particular, each participant's ah close read and interpreted poems. ah and then like And then each participant was randomly assigned to having no AI or a single AI interpretation.

Speaker: or multiple AI interpretation but when when they wrote their interpretation. And then basically for each participant, they were asked to interpret three poems in total. And then after their interpretation of each poem, like we also asked them about their own experience on that.

Speaker: on or that interpretation. So based on that, ah like this experiment examined both people's um performance of interpretation. So that is the output there, as well ah like ah and as well as their pleasure of close reading. So that is more about on their that's more about their experience in the close reading process.

Speaker: So, ah so with that said, on one side, we have the ah interpretation performance, which is the output. And the other hand, we also measure people's like close reading experience, which is, so like in yes in terms of like, did they enjoy it? Did they appreciate it? Did like some sort of confidence, but you didn't. So unlike your other study, right? You didn't judge like,

Speaker: What details do they remember about the poem? Whether they could like analyze the poem independently on their own in an interview setting without the AI, right? like you did measure like Basically, you didn't kind of do a measure where it's like you tested how well they understood the poem independently.

Speaker: of Is that right? Oh, ah it's like... ah i think ah so I think that would actually be the condition where the participants ah were assigned to having no AI. so and So that's the condition where they need to write their interpretation of the poems independently. Right, right, right. So it's like of the 400, there was one group that had no AI interpretation. But you didn't... just but like I'm just thinking how in your other um research, you you do some stuff where it's like some people will use AI to some degree to do a critical thinking task.

Speaker: um But then they're also kind of tested after where they don't have the AI and you kind of see how well... they grasp what they did, right? Isn't that correct?

Speaker: ah What this experiment did is that like this, a so it's like this already made single or multiple AI interpretation, is provided to the participant so that the put are so that the participant didn't need to instruct the AI themselves, oh which is- Okay. Yeah, so ah which is though like which is a way to reduce the variance that might be introduced based on facebook how people use and interact with the AI? So it's more like a design choice for this close reading experiment.

Speaker: That's really interesting to me because in some other experiments that I've read, not not your experiments, but other experiments on usage of LLMs, they have, I don't know what the technical term is, but something like laissez-faire, like they just give them access to AI and participants do whatever they want with them. In your case, these are prompts that are provided or not even provided, right? It's actually all they see is the output of the LLM without. Okay, cool. So...

Speaker: And you're saying relative to ah participants that had no AI usage, those that had one AI interpretation ah performed better. ah And at least if they were inexperienced readers, they enjoyed it more. is that Did I get all that right?

Speaker: Yes, that's right. but Wasn't there like a trade-off though? Can you go into that? I thought it was like the case that... oh I mean, i don't know if this is maybe not inconsistent with what Roberta just said, but I thought like one element, one finding was like this interesting thing where the more people relied on AI, the more, the less pleasurable they found the whole experience, right? it wasn't that true. Like they valued, they didn't find as much enjoyment and pleasure in the close reading when they heavily relied on the AI interpretation, right? Isn't that, or am I wrong?

Speaker: So basically there are two perspectives to look at the experiment results. So but let me ah unpack this ah one by one. So on one side, if we only look at the directionality effects of those conditions, we found that ah overall um by having a single AI interpretation or by having multiple AI interpretation, participants' interpretation performance were improved.

Speaker: But it is a more nuanced case on for their experience. It's like we found that ah only by having one AI interpretation, like participants' enjoyment were improved. And then like if they were exposed to multiple AI interpretation, like there's no such effect.

Speaker: And then if we look at these experiment results from another perspective, especially i look into how people made use of those AI interpretation ah provided to them.

Speaker: And then like, especially I found that ah when people's a written interpretation high textual overlap with the AI interpretation, they would actually show lower enjoyment in that process.

Speaker: So I think it's fair to say that people just ah rely on the AI interpretation in a simple way, such as just copy and pasting the AI interpretation. Like it is hard to enjoy the undiscovering process so that this close reading process would be less enjoyable for them, if that makes sense.

Speaker: Yeah, copy and pasting is less fun than yeah thinking and figuring out what a poem means and writing about it or whatever. yesam Yeah, I also saw in the in your paper that in their feedback, maybe it was the ones that copied and pasted a lot, or maybe it was just the ones that had access to several AI interpretations. But in their feedback, some participants said that basically the AI output was so good that they felt intimidated, right? Like there's almost no room for ah for for you know human input at that point or for them to personally add anything to that. And it reminds me of like ah Lisa Doll who gave up on the game of Go when ah you know ah an AI beat him. right And he said, well, now there's something, this isn't interesting to me anymore. right I'm not sure that's exactly what he said, but if the AI output is too good and you're exposed to too much of it, to three in this case, that's enough to kind of take away the the enjoyment

Speaker: Yes, exactly. Especially for the case of multiple AI interpretation. Although on one hand, we know that one poem may have a lot of reasonable way to understand its meaning. But on the other hand, being exposed to multiple AI interpretation like there's really like minimal room for human interpretation.

Speaker: so like in the case of being exposed to AI interpretation, like ah people would lean towards just offloading this ah entire interpretation word to the AI and then not doing their own thinking on and of the poem. And then like, that's what makes this close reading process less enjoyable for them.

Speaker: One last thing I want to ask about this experiment before moving on to the next one. You did write that, so I guess let me back up a little bit. um We interviewed a guy named Michael Gerlich, who you cited in your research. And one of his big things is that he calls it social bifurcation, where he says that Some people will become overly dependent. Actually, the majority of people will become overly dependent on artificial intelligence. And there will be a resilient minority that will kind of retain their interpretive autonomy and do the cognitive work for themselves. Mm-hmm. But you seem to be reporting a finding that might go counter to that because you write that ah readers i seem to have like a natural, healthy resistance to letting AI take over their cognitive work, their interpretations.

Speaker: So you want to ask you about that. You say two out of five participants denied using AI altogether. So, I mean, can you tell us about that? Did you check that they didn't use AI, all that?

Speaker: Oh, ah ah yes, I would love to talk more about that. So, um yes, ah I think that percentage, about 40%, like, is a bit striking to me. So, in particular, ah it's like 45% the single AI condition and then in

Speaker: interpretation condition, like reported not using AI, even though the AI interpretation was right there. ah And then like I also look at whether that self-reported AI use aligned with their behavior, in particular um among those who self-reported and not using ai it's like more than 90% of them actually had no copy and paste behavior at all.

Speaker: But I think the nuance here is that in their written interpretation of the poem, it's like they still show some textual overlap with the AI. So I'm not saying that they were misunderstanding or like lying about their AI use. So I think it's more like people's own perception, like I didn't use AI, it's really their own perceived non-AI use. So in that case, they may have actually absorbed or digested something from those AI interpretations without perceiving it ah ah as their own AI use. So

Speaker: I think it's more about how people perceive their AI use, ah like which is not a clean behavioral fact, but like it is kind of fascinating itself because it's suggesting a way that AI can influence what we think ah like ah ambiently like bypassing our conscious awareness, if that makes sense. There's some philosophers who use social psychology to talk about how ah it seems to be the case that the environment has an effect on our actions without us really noticing it. And we attribute it to ourselves when, you know, under behavioral constraints, you can sort of get people to do one thing instead of another based on the situation. So it seems like... um

Speaker: It doesn't have to just be rooms. It can even be a computer screen with ah an AI, you know, output or interpretation on the side. And that's enough for you to say, okay, I i didn't look at it. I did it on my own, even though you it did seem to influence you. So I think Jayian has some really good social psychological support for that. That should be a paper in and of itself. That's fascinating.

Speaker: So could I real quick ask, you know, I'm just imagining someone who has influence in educational policy or something reading your abstract and they might be like, oh, single AI interpretation boosted both performance and pleasure.

Speaker: So that means um by having AI more in the mix when people are reading poetry, the individual becomes a better close reader because their performance was better. The person will personally do more interpretive reasoning.

Speaker: They're going to develop more transferable close reading skills. But, but I mean, isn't it the case that like your point here in terms of performance was more like evaluating the ultimate output.

Speaker: So when it comes to the ultimate output, artifact, you know, the interpretation of the, of the poem that the person sort of like submitted.

Speaker: It's the case that, you know, when it comes to like, like identifying different features of the poem or like, quality of the written interpretation or like the quality of the writing, that sort of thing that proved, um, when there was a single, single AI interpretation boosted performance in that sense. But it's, so I guess I'm just saying that you didn't necessarily like test whether the person like independently became better, closer, close readers, like on their own.

Speaker: You see what I'm saying here? um Yes, I think for this experiment, ah yes, like I didn't take into account like what this AI-assisted interpretation ah leave people to, which is some something I think it will be definitely ah worth investigating like as a next step. And then like for this experiment,

Speaker: specifically because like the population i focus on is like lay people, is like ordinary people, especially for those ah lay people who may read on text for entertainment. So it's like in a way for, on on ah so here I think the thing is ah close reading performance and then like the boost in their close reading or like ineptation ability is not like a central ah concern for this type of close reading purpose for entertainment. So it's actually more about like the experience and they like more about this appointment ah night right and then more about more about, yeah, like yeah like it's more about this enjoyable experience that people ah ah that people can derive from

Speaker: from that exactly right like like So for example, like a high school teacher shouldn't necessarily read this article and think, oh, okay, so if I want my students to be good at close reading themselves, I need to start giving them more AI-generated interpretations of whatever text they're reading. like That's not necessarily the right conclusion to draw from this, right?

Speaker: So I think that would actually be more and nuanced because um because because i based on my views and experience, I think like any experimental evidence ah should be interpreted be understood its implication in a way that fits the ah actual context. So and for example, for educational context, I think like I think, like as you mentioned, things like maybe high school teacher like ah ah can see this treatment condition of a single AI imputation as this modest exposure to AI assistant. So I think in educational context, the actual implication would be it's like to assist students learning maybe like a modest amount of exposure to AI assistant

Speaker: could be helpful. And then like, I think the case of multiple interpretation meaningful for the educational context is that say like if for learning purpose, if this room of interpretation is consumed by ai it's like, it's like there's not much room for the students to do it themselves, which is something that we probably don't want for learning purpose, if that makes sense.

Speaker: Can I ah put put out a description of an experiment and you tell me if that's sort of what we would need in order to kind of give advice to teachers? Let me just lay this out, see see what you think.

Speaker: So what we need, I think what would be a nice follow-up experiment is something like, To harp on something that I said earlier, it it wouldn't be laissez-faire use of AI. So we need an experiment where participants are given the prompt um designed by their teacher, presumably.

Speaker: And we have the intervention and we evaluate that, sure. But then there's also post-AI intervention AI independent um measurement to see if if there is a difference without machine aid ah for those people? is it Would that be able to kind of really give us some actionable ah you know tips for the educational sphere? Yeah.

Speaker: Yes, I think like i so i think like basically by having a delayed post intervention, like this AI independent test, like it will give us some evidence about like on the delayed impact of of using AI for something. Yes, I think that makes sense.

Speaker: Okay, let's move into your second experiment going to talk about today. It's a critical thinking task. And there's essentially ah two time conditions. Some people get enough time for the task, like 30 minutes, I guess.

Speaker: And some people get insufficient time for the task. I think it was 15 minutes. And then there's three... Ten minutes. I think was ten. Ten minutes? Okay. And there's four AI conditions.

Speaker: One group has no AI. One group has AI throughout the entire time. One has AI in the beginning, and one has AI access at the end.

Speaker: ah Maybe I'll let you... Only at the beginning or only at the end, right? Right, right, yeah. So maybe you can kind of tell us if did I get all that right, first of all, and then also kind of tell us so why you chose to do that kind of study and what you were trying to discover.

Speaker: ah Yeah, I think you have just described the condition design. ah So yeah, like, so basically like started from the motivation. So, ah so I think I did this kind of research because I ah like,

Speaker: Because I want to answer like this question around like what is AI doing to critical thinking. So I think here like the core question is basically say that if we upload tasks that humans have traditionally done to AI, like what would happen to the thinking that we do ourselves? Does it lead to like better... thinking outcome and then if it does does it mean our thinking is actually getting better and then like i think like overall old trend is that it is shifting our um cognitive work from implementation to supervision uh uh specifically now with ai like oh like we do ah

Speaker: It's like we do less of the implementation and then like more and then like more of the checking and then like deciding whether the AI output is good to use.

Speaker: So here the question i want to investigate is that like whether people have or if people actually have that capacity to think critically with AI.

Speaker: So that's why and i want to study AI's impact on critical thinking, which is like this capacity at the center of such conversations. So this...

Speaker: so like ah so like basically this um experiment is about AI's impact on people's performance at a critical thinking task. And then when we really get into this question, time is a factor that cannot be overlooked because it is related to people's cognitive mode. specifically, like ah it is well established that ah if people work under ah time pressure, on they would lean towards fast automatic processing ah rather than a slow undeliberative reasoning on that critical thinking would require. And then like such thinking mode under time pressure would extend to the occasion when people do tasks that require

Speaker: critical thinking with AM. ah So that's why on one hand, this experiment takes this time availability for task completion into account. And then like, and then like, as you just described, on and I basically tested what differences it would make if people have sufficient or insufficient time for completion.

Speaker: task completion and then like on top of that on time availability for task completion and I tested different ah timings of AI access like in a way to steer on if the AI is used throughout a task like used only for those preliminary early stage cognitive activity or used only for ah late stage wrapping up like activity or not used at all so like In this way, like this experiment could give us ah like a more un holistic understanding of AI's impact on people's critical thinking performance, if that makes sense.

Speaker: Yeah. And one thing I liked about what you guys measured here was like, so, you so yeah, you have, I think it was like 393 participants or something. And, you know, you're giving them source documents related. i think the case was related to like water contamination and all of those people are producing an essay at the end of it. Right. Yeah.

Speaker: and But you didn't just evaluate like the quality of the output. It wasn't just like, you know how many valid arguments are in this output? It wasn't just, you know how many references does the document and does the output include to the documents?

Speaker: You also... measured like the recall of the participants themselves. So it's like you ask the participants, Hey, like, you know, can you remember this or that from the document? You know, can you kind of like measure their own comprehension?

Speaker: hmm. Anyway, so you kind of get both a picture of of the impact of AI or no AI on um the final output, the product, the artifact, the essay. but then you also get a sense for what kind of you know impact does engaging with AI have on the person's like independent, personal, critical thinking ability, their comprehension, and so forth.

Speaker: Yes. So as you just mentioned, so here it is critical thinking tasks that I employ. So ah it is like a validated ah instrument from educational psychology that is designed to measure people's ah critical thinking abilities. And here, this final output, which is the argumented essay here. So it is in a way to capture like this final outcome of this critical thinking process and then like a critical thinking as a comprehensive a higher order on cognitive activity it's like it encompasses like different ah intermediate cognitive activity which includes comprehending the information and then are like ah internalizing the information and then like on evaluating like the source of information and then like on it

Speaker: evaluating and then weighing the trade-off between different perspectives derived from that information. and Finally, it involves ah like synthesizing those on different pieces information. So as the final outcome, like this argumentative essay, like a measure, this ah final output of the critical thinking process, in addition to that final output, and I also measure people's ah record, comprehension, and then explicitly their evaluation of those source documents, like in a way to ah measure not only the output, but also those stages.

Speaker: Yeah. And just so like real quick, like one result you found that kind of draws this out is that you found like if someone had insufficient time, so they're only given 10 minutes to do a task that really should take, that would normally take like 20 minutes or something like that. They're only given 10 minutes and they're given an LL from the beginning, from the get go.

Speaker: You know, their essay performance might be really high in the sense of like the final artifact that they submit ends up being pretty good.

Speaker: But that doesn't necessarily mean that they get what they submitted or like actually understand. Right. So that, so that, the so their scores on like comprehension and all that stuff might've been lower. Right.

Speaker: Even if their essay performance was high, right. um ah in those situations where they had insufficient time and were able to use the LLM from the start, right?

Speaker: Yes, exactly. Because when people work under time pressure, so here it's just 10 minutes, like for a test would normally require about 20 minutes. So it was almost impossible for them to read those documents themselves. So for those participants who have an AI to use from the start, it's like they are just using AI to come up with an as essay ah instantly and to have the test done. It's like they are not really digging into those documents, ah which could explain ah like which could explain why their record as well as comprehension performance was low.

Speaker: On the flip side, i want to, I guess we can't go over all your findings, but another interesting finding that we should highlight is that when you are given sufficient time, ah first of all, and you have ah You basically do your own independent work first and then get access to an LLM.

Speaker: ah You show all kinds of improved ah performance on the metrics, right? Good essay, but also greater recall, right? Did I get that right? well Is it improved in comparison to the person...

Speaker: who doesn't use ai at all. Yeah, maybe. because that's yeah Actually, could you do what? So could we compare these two? What's what's like the comparison between the one Roberto just mentioned? So you're given sufficient time.

Speaker: And then Roberto, you said you're given the LLM late, yeah right? How does that compare to someone who's given sufficient time and doesn't have the LLM at all?

Speaker: Yeah, so basically, when we put people under sufficient time and then if we compare those who have like late AI access to those who have no AI to use at all,

Speaker: it's like they actually show ah very close assay performance. but like on the other hand ah ah But on the other hand, it is interesting that those with late AI access, on they actually show ah reduced mindset bias ah compared to those having no AI to use at all. So that is the main difference between like having late and no AI access when people have...

Speaker: and so just So my side bias, that refers to like, so like let's say let's say like someone writes an essay on whether God exists and they give five arguments for the existence of God and zero arguments for That God doesn't exist. That has high my side bias because it's all arguments pro.

Speaker: Yes. Yes. Whereas if they gave like three and three, that would have ah low my side bias. Right. And so the finding is that if somebody has a lot of time or whatever, sufficient time, 30 minutes.

Speaker: and they don't use LLMs, they are going to have higher my side bias. They're going to be more partial to one viewpoint. They're going have less arguments contra than somebody who who uses an lm LLM late, right? is that basically what you're saying?

Speaker: Yes, that's the case I found in this experiment, like which is reasonable because imagine and if someone is doing like a solo ah deep dive and then this entrenched ah one-sided reasoning ah ah and it's actually like a natural tendency. like Right. Yeah, like if someone is working like completely independently,

Speaker: But like when they have like AI as a second role, it can potentially reduce that mindset bias um by introducing some counter arguments into their essay.

Speaker: Yeah, that totally makes sense. Just to go back to another thing you said, though, in terms of the final output, I found i found this like really interesting. So if you're just looking at the quality of the final essay, you're just looking at, hey, is this essay that I'm reading right in my hand, does it have a lot of valid arguments? Does it blah, blah, blah?

Speaker: Mm-hmm. The person with no AI and sufficient time does basically just as well as the person with sufficient time and uses AI late. and Right? So i just found that interesting because it's like it's not as though like, oh, if you want a really good product, if you want a really good essay, like better use that AI at some point. like No, actually, no AI gets you about the same result, sounds like.

Speaker: Yeah, like in a way, ah yeah, like i found that fascinating too, but like that is also a conditional. You got to have a lot of time ah for you to do that independent work. And then like this is like especially a task that people can have like independent good performance.

Speaker: So that is the case. ah So I think here the point on this experiment is making it's actually like when people have a really sufficient time for a test, oh which is also a test on dad that they could independently complete if they have really sufficient time. It's like by having an AI to a system, on that doesn't make much difference for the final output and if it's about like coming up with an argumentative essay about this issue. Yeah, so I think that's the point. of

Speaker: But like here we can see like there are some unique benefits by on having that AI a assistant can make like which is like this reduction of my side bias.

Speaker: I think that's the point. Another thing that I want to ah bring back into the picture is that this is laissez-faire usage of LLMs. These participants are are allowed to use them however they wanted. And at the end of your paper, you even kind of talk about all the different uses of that.

Speaker: But I wonder that you know if the the benefit of using large language models in reducing my side bias, I'm wondering if it could even be, you know, expanded to have other benefits had it been, i don't know what to call it, like ah like the structured prompts that were provided as in your last study.

Speaker: i can imagine like an English teacher or, i mean, I don't know, I guess an expert in the domain, whoever that might be, if they craft ah prompts,

Speaker: to give to participants and they get to, you know, kind of choose and whatever based off that. I wonder if that would have a more pronounced effect, you know, on not only reducing my side bias, but maybe even boosting performance overall, almost having an AI as, as a, not only a sparring partner, reducing your my side bias, but also maybe scaffolding the process or something like that. Does that make sense?

Speaker: Yeah, I think that makes total sense. For example, if the design goal of of that AI system is to like reduce people's mindset bias, or like in another term, ah if the design goal of the AI system is to ah is to have people consider more diverse perspective, and then also in a way to have them resist resist their like stereotypical mental shortcut,

Speaker: I think by having like a scarf folded system would definitely help if that is the design goal. And then like I think it would be interesting see like more AI system design in that way, and especially for educational context. Yes.

Speaker: Could also just ask real quick, um was there any difference between... Okay, so for people that's sufficient time... When we compare the group who had AI only late versus the group that had no AI, um was there any differences in terms of like recall, evaluation, comprehension? So these things are, my understanding, it's like this is kind of evaluating whether the person independently, you know, but was able to recall things, evaluate, comprehend things after the fact.

Speaker: um Yeah. so was there any difference there? Like do the people who had ai late, did they have better recall after the fact about anything like that? Actually, for for metrics, including like people's ah essay performance, record and comprehension, and evaluation performance, it's like for those aspects of performance, on participant having late AI access and no AI access, it's like they actually show pretty similar.

Speaker: performance. Yeah. Okay. So it's just like with the SA output, like the SA performance was basically the same. And you're saying for the for like the recall, um comprehension, evaluation, basically the same. So that's really interesting. So the only, yeah. So like you said earlier, the only real separation in those groups has to do with the my side.

Speaker: Yes. Yes. That is the main difference between late and no a access when people have sufficient time. Gotcha. Yeah.

Speaker: I'm looking here at the time. I'm wondering if we can ask you one ah big meta question to kind of ah i take us home here. Yeah, sure.

Speaker: Let's just say that we want some advice for how to use large language models um if and when they can benefit our tasks without ah degrading cognitive capacities, or becoming dependent on them, but also without um impeding our learning so that we can be good at that task independently later on. it seems like one thing that we have to do is work independently

Speaker: and then maybe later use ah large language models. But do you have any other implications of your work for how it is that we should use LLMs profitably?

Speaker: ah Yeah, I think that's a great question. also ah Also really a comprehensive one. so So I think like basically ah broadly speaking, like now that AI is basically ah infusing in almost every aspect in our life and work, and then like it is accelerating how we can get things done.

Speaker: For example, if we're looking at the insights from this experiment, when we talk about time pressure, ah which is like a very common ah ah case in real world practices. So I think the implication here is like, if we are working under time pressure, like I wouldn't say it would be the beneficial case if we use AI from the start.

Speaker: So it's more like it would be nice to like get a sense of like what AI is doing ah if we use AI under time pressure.

Speaker: And then like based on the experiment result, like on we did find working independently first um with AI assistant on top is the best condition of performance overall.

Speaker: That could potentially be understood as the most ideal case of using AI, ah oh which is, and as you also described, which is independent work before AI use.

Speaker: But like in reality, ah in many cases, we don't necessarily get to do that independent work before AI use, ah especially when we are working towards some deadline ah under some ah real world time constraint.

Speaker: But on the other hand, it's fair to say like it is important to raise our awareness about AI's impact on our thought, especially with consideration with those real-world time factors.

Speaker: Well, we've been talking to Jai-Yun Zi, and Jai-Yun is a PhD student studying computer science at the University of Chicago. Thank you so much for your time. Thank you. This is fun.

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Speaker

Recommended