I wonder if this will change the ballgame (and the trawling web page content):
Data is not copyrightable and it’s not plagiarism simply to report it. If I say that the “United States has 50 states” it’s neither plagiarism nor a copyright violation despite the fact that I’m sure that many many people have said that data (and likely in that exact form). Plagiarism is the exact copying of a substantial amount of another work, presented as the author’s own. My understanding of AIs is that they don’t use the same phrasing as the sources they train on* but instead paraphrase, as a human author might.
*I think there were cases were the AI spit out exact paragraphs but that’s not the common result.
It need not be exact. Wikipedia addresses this (as well as the distinction between plagiarism and copyright infringement).
[…] the practice of plagiarizing by using sufficient word substitutions to elude detection software, known as Rogeting, has rapidly evolved. “Rogeting” is an informal neologism created to describe the act of modifying a published source by substituting synonyms […]
Fair enough. If a student takes something from a single source and changes only a few words, it’s considered plagiarism. Or if they take an idea that exists only in a single source and doesn’t cite that source, even if they write it in their own words, it’s plagiarism. But if a student reads ten books (about, say, the Civil War) and then writes an explanation of that war in their own words (ie paraphrasing)? That’s not plagiarism. The latter seems to be much more what AI is doing, but the criticisms I’m seeing aren’t making any distinction whatsoever.
If a student does so w/o citing those sources, it’s plagiarism. Same for a scholar. The student/scholar may not need to cite specific passages, but to be academically honest, would provide a list of works consulted so that the specific material the writer is drawing on can be checked. Besides: you’re gonna find that the ten sources will not agree on everything, and that the writer will have made choices about which source to follow for a given statement. That choice needs to be cited - by student or AI.
It’s really not. If this is a graded assignment in an academic class, it might be reason to grade them down for not having the proper scholarly apparatus but it’s not something to report to the academic integrity folks.
The Civil War started in 1861, between the northern and southern regions of the United States, over the issue of slavery. It lasted four years and inflicted enormous damage on all regions of the country. It remains the most important war in American history.
I’m a scholar and that summary is based on sustained reading of a lot of Civil War histories over the years. Did I just plagiarize?
(Your knowledge of what plagiarism is almost certainly came from reading a definition. You didn’t cite it. Did you plagiarize?)
My personal concerns about LLM/gAI scraping are more about privacy and control than whether or not it constitutes plagiarism, no matter how one defines “plagiarism”. I also think it is important to recognize that while everything available online is indeed data, not all data is free of copyright and creators’ rights.
An editor would probably take a scholar to task, but as a scholar, you know that. You’re right that what is considered plagiarism is contextual; it even varies across disciplines. I used to have long lunch discussions with my colleagues in business and business ethics about this matter. In our discussions, we rarely cited - but we did when we were being called on a claim, and we cited when we wanted to move the discussion forward. There was the occasional, "You should read this … " with a copy of an article or a book. Citation serves that purpose: to extend the creation of knowledge.
My scholarly context is composition research at one end and literary theory at the other.
I see two statements, maybe three, in your summary that would deserve citation if you placed the summary in a scholarly context: you can find them yourself. It’s more a recitation of common knowledge as taught in schools than a summary - which says a lot about what we consider knowledge.
But your summary illustrates two issues that apply to AI: a) You didn’t ‘explain’ the Civil War, as your example in your earlier post called for; you summarized a set of sources on the Civil War. Some of your summary is made up of statements of facts; some is made up of arguable claims. Both deserve citations in more formal contexts; I’d argue that the sleeping claims demand citation, even in an informal context. It’s very like AI: AI generates summaries but not sound explanations. And, importantly, the confidence in AI summaries often masks the difference between statement of fact and claims. You may have even used AI to generate the summary. And b) You didn’t really say anything about the Civil War that adds to general knowledge: nothing new, no explanation. That suggests a limit to AI’s range: it churns out stuff but doesn’t add much that is new. It doesn’t add to a larger body of knowledge.
My knowledge of plagiarism stems from 35 years of teaching academic writing at university level, and writing it. I’ve read a few definitions. I’ve written a few, too. Did I plagiarize in my consideration? Nope. I stole, or maybe I borrowed - and you surely know the (disputed) source of that allusion. [citation needed]
Thanks for the exchange, but I’ve moved us off-topic.
Yes – which is not plagiarism and seems to me to be largely what AI bots are doing: reciting common knowledge.
No, I explained the Civil War, using the knowledge I had gained from a wide range of sources. No one inherently knows about the Civil War – the knowledge has to come from somewhere. The definition of plagiarism I see people implying in this thread comes dangerously close to making any discussion of knowledge that doesn’t come from first hand experience into plagiarism.
And if it’s only repeating common knowledge, then it’s not plagiarism.
I don’t think you have, actually. I think a key problem here is that when it gets labelled plagiarism or copyright violation, the discussion gets mired in a legalistic debate over those concepts rather than focusing on the gross misuse of other people’s work to create things like ChatGPT. The latter is the issue, not whether certain formal definitions apply (as @Halfsmoke noted)