Jump to content

Wikipedia:Village pump (WMF)

Add topic
From Wikipedia, the free encyclopedia
Latest comment: 5 hours ago by The Other Karma in topic Source Verification Suggestion
 Policy Technical Proposals Idea lab WMF Miscellaneous 
The Wikimedia Foundation (WMF) section of the village pump is a community-managed page. Editors or Wikimedia Foundation staff may post and discuss information, proposals, feedback requests, or other matters of significance to both the community and the Foundation. It is intended to aid communication, understanding, and coordination between the community and the foundation, though Wikimedia Foundation currently does not consider this page to be a communication venue.

Threads may be automatically archived after 14 days of inactivity.

Behaviour on this page: This page is for engaging with and discussing the Wikimedia Foundation. Editors commenting here are required to act with appropriate decorum. While grievances, complaints, or criticism of the foundation are frequently posted here, you are expected to present them without being rude or hostile. Comments that are uncivil may be removed without warning. Personal attacks against other users, including employees of the Wikimedia Foundation, will be met with sanctions.

« Archives, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17

AI-generated edit suggestions

[edit]

When Simple Summaries rolled out, the overwhelming response from the community was that we do not want AI features. One year later, we are now getting more AI features, e.g., AI-generated edit suggestions. This has been in the works for about 2 months and is due to be announced any time now, so you heard it here first. This is going to be long, sorry, there is a lot of ground to cover.

First: These appear to be the suggestions, because based on how hard that URL was to dig up I assume the announcement wasn't going to link to them in one place. (It's unclear which of these lists if any is the actual list used in production, but there do not seem to be major quality differences based on spot-checking several of them, the newer lists are not obviously better and the larger lists are not obviously spottier). There are three broad problems here, besides the fact that adding new AI features is the opposite of what people have asked for:

  1. Besides the most obvious low-hanging fruit (typos), the suggestions do not contain any concrete fixes. I am pretty sure this is due to not wanting to create a vector for adding AI-generated text (from the study: We don’t have any plans nor desires to use models to create edits directly), and I agree with that. The problem is that, based on the kinds of suggested edits we've seen in the past (e.g., Newcomer Tasks), people will resolve this ambiguity by using AI anyway.
  2. To that point, it is unclear who this is for. A lot of the suggestions are obvious low-hanging fruit that competent copy editors would have noticed on their own anyway. People who are not fluent in English will not be able to do much with a suggestion like "Rephrase the sentence to present the information in a neutral tone and qualify the superlative with a time reference." Newcomers who might have been overwhelmed will probably be even more overwhelmed because of the lack of direction.
  3. Several of the suggestions are just bad. Sorry, but there's no kinder way to put that: they are bad suggestions that encourage bad edits. This is just a spot check and is incomplete, but I've identified a few common categories of bad suggestions (drawn from multiple lists at the link):

Political/geopolitical nightmares: I strongly suspect the geographical name suggestions are going to or have run into some geopolitical snarl at some point. But the NPOV suggestions already have, and demonstrate the reason why using AI to "correct" non-neutral point of tone is the opposite of "low-risk" (as described in the writeup) and generally a bad idea.

  • FreedomWorks: This article about a conservative organization contain(ed?) the sentence "During the 2020 election campaign, FreedomWorks pushed false and misleading claims about mail-in-voting, targeting ad campaigns on swing states with high concentrations of minority voters." It is cited to a Washington Post article describing clearly false and misleading claims. The AI's suggestion, however, is The original wording presents FreedomWorks' actions as definitively false and misleading, which is non‑neutral. It should be rephrased to a neutral description of the disputed nature of the claims.
  • Yishuv: The original article contains the following sentence: "The League of Nations codified support for the eventual 'establishment in Palestine of a national home for the Jewish people' into the foundational document of the British Mandate in Palestine, thereby facilitating what detractors later regarded not as aliyah, or the flight of refugees from Nazi and Fascist atrocities, but as the Zionist colonization of Palestine." Obviously that's not perfect, but this suggestion seems to be a misinterpretation: The passage uses loaded language that frames the British Mandate support for a Jewish national home as ""Zionist colonization"", reflecting a partisan perspective. It should be rephrased to present the differing interpretations without endorsing one.
  • The LLM is oversensitive to even the most factual descriptions of political views or affiliations. Example: DeAndrea G. Benjamin, re: the sentence "During her confirmation hearing, Republican senators questioned her decisions granting bond and early release of defendants": The phrase ""Republican senators"" introduces a partisan label that is unnecessary for a factual description of the hearing. The senators who raised these questions are factually members of the Republican Party, all three sources frame it as such, and the partisan breakdown is an inherent fact of the situation.

Suggestions to introduce factual errors:

  • Frost Children, regarding a sentence about the album called Smile! :D: The sentence contains an emoticon, which is non‑neutral and informal. If someone followed this suggestion they would turn a correct statement into a wrong one.
  • 1797 in Denmark: The word "pyusician" is a misspelling; it should be corrected to "musician". The actual correct spelling here would be "physician".

Parsing errors producing nonsense: This is the same problem Simple Summaries had. The text parsing breaks in many ways, and the LLM generates suggestions based on the broken version.

  • The LLM has trouble with wikitables, and will often directly reference a JSON snippet it received, resulting in bizarre suggestions like Remove the nonsensical JSON list and replace it with a brief, readable statement or omit it entirely. (Markazi Jamiat Ahle Hadith)
  • The tool seems to assume that excerpts are one full sentence long, even when they're not. This results in several suggestions to "break up" sentences that already are. This happens a lot, but an illustrative example is Santa Fe Place (the "Stores" paragraph), where the actual prose problem is the opposite as Overly long sentence with many clauses; needs to be broken up for simplicity): the sentences are choppy and some could be combined. (also there's an obvious comma splice the AI fails to mention)
  • Sometimes titles, sidebars, etc. get interpreted as article text: 2019–20 Philadelphia 76ers season: The lead repeats ""NBA professional basketball team season"" twice, creating redundancy. Obviously, it does not; the culprit is that the sentence the LLM interpreted was "NBA professional basketball team season NBA professional basketball team season The 2019–20 Philadelphia 76ers season was the 71st season of the franchise in the National Basketball Association (NBA)." (See also Boadicea Haranguing the Britons, where it does this for the template)
  • This also happens with templates, as in Chen Lijun (actress): Redundant repetition of the subject’s occupation; the phrase “Chinese” appears twice. The first "Chinese" comes from the "lang-zh" template.
  • Cross-wiki links get it too, as in Switzerland in the Eurovision Song Contest 1966: Extraneous ""[it]"" markup after Mascia Cantoni; it is an artifact of extraction and should be deleted.
  • So do blockquotes, as in Nacht und Nebel: The passage contains a run‑on sentence and an incomplete citation phrase ""According to historian Wolfgang Sofsky:"" that leaves the reader expecting a quotation.
  • So do stub templates: The stub template line includes an unnecessary ""vte"" fragment and could be phrased more cleanly. (the "vte" shows up a lot) or Remove the redundant stub messages that appear as ordinary text at the end of the article.
  • Some suggestions are already fixed. For instance, Bomb-making instructions on the Internet was (very obviously) vandalized in Special:Diff/1358132635, and the vandalism got reverted by ClueBot basically immediately. The suggestion nevertheless refers to the vandalized version (Remove the non‑encyclopedic, opinionated rant). I don't know how the LLM got hold of a revision that existed for only a few seconds.

Suggestions to conceal the symptoms of a larger problem:

  • As seen above in the bomb-making example, the LLM doesn't seem to know about vandalism and will never describe it as such, regardless of how obvious it is. This seems likely if not certain to encourage someone to "fix" the tone of vandalism without addressing the actual claim, which is how we get years-long hoaxes.
  • One thing that happens fairly frequently is that the suggestion feature will flag an article that, in context, is clearly AI-generated. (Examples: Jeremy Coller, Vimbuza). It never picks up on this and suggests minor tweaks to wording that would just put a band-aid on the issue. (The csv is actually fairly useful for this, but only in full searchable list form.)

LLM-specific tics: Obviously the suggestion text itself are AISIGNS overload but what I mean here is that LLM edit suggestions/summaries have some consistent quirks. I haven't done an in-depth look for them but two known ones do show up frequently:

  • For some reason LLMs have a fixation on "superlatives" being inherently non-neutral, when sometimes they are just true. I don't know where this comes from -- WP:NPOV doesn't mention anything about superlatives -- but it shows up all the time in LLM-generated revision suggestions, and also all the time here. For instance, in List of Hercules: The Legendary Journeys and Xena: Warrior Princess characters: The description of Hercules is too lengthy and includes biased phrasing such as ""strongest man in the world"" [...] Or on Trinidad, California, a sentence starting "On December 31, 1914, the largest recorded ocean wave ever to hit the United States West Coast" (in a paragraph with 5 citations) is criticized with The sentence makes an unqualified superlative claim about the wave’s size, which could be seen as non‑neutral. Adding an attribution phrase mitigates this.
  • LLMs will over-justify anything. The following reads like a parody but it is actually a suggestion from these: The word 'ithe' is a typo; it should be 'the'. Fixing this error improves the sentence's correctness.

I don't really know what to say at this point. Based on the convoluted phabricator issue snarl it seems that there was some human review done at some point, but nevertheless it took me only about ~1-2 hours to find the above issues, and that was only a spot check. This seems like a reasonable amount of human QA to expect before pushing an editorial feature to prod. It also seems reasonable to expect a full audit of potentially controversial subject matter (by "full audit" here I mean even just the results of CTRL-F "Israel", "Palestine", etc.), and ideally a review by people who are familiar with AI-generated revision suggestions and the general areas in which they get things wrong. (None of this deviates much from the thousands of similar justifications I've seen in AI-generated edit summaries.)

I also think the whole premise is just flawed. LLM and machine learning tools are not only going to have high rates of false positives, but false positives that take time to evaluate. (For instance, there are various LLM tools to scan articles/edit summaries for possible issues, but the point is that they generate lists to be manually reviewed later.) They do not scale to a scenario like automated edit suggestions, where the assumption is that the suggestions are pre-vetted and can be evaluated quickly. Gnomingstuff (talk) 17:54, 15 August 2026 (UTC)Reply

@Gnomingstuff I get your frustration, and agree with some of your points, but I still think this is worth a try. I will often ask the LLM-bots to proofread my articles. Some of the suggestions are obviously good (mostly low-level stuff like spelling, repeated words, etc). That's the kind of stuff that once you've read your own writing 100 times, you read right past and don't notice, so I find it an invaluable service. The higher-level suggestions (tone, phrasing, flow) I'm much more likely to reject, but I accept them often enough that it's worth doing. But that's really no different from when I'm working with a human reviewer. I'll often push back and say "Nah, I think the way I've got it now is fine".
So I think the trick here is to figure out how to educate people that these really are just suggestions and they need to apply their human judgement about whether to accept them or not. You are correct that for new editors, that may be problematic. Still, I think this is something worth trying as long as we monitor how well it's working out and be willing to pull the plug if it turns out to not be useful.
LLMs are just the most recent technology step between scribes writing on clay tablets and where we are today. We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" Or we can experiment with them (inevitably with some failures) and learn how to take the best advantage of them to improve our product. I vote for the latter. RoySmith (talk) 18:17, 15 August 2026 (UTC)Reply
Please don't ping me to a discussion that I started less than an hour ago and am clearly aware of.
I don't think that We can dig in our heels and chant "LLMs bad, down with AI, all power to the humans!" is a fair assessment of something that I spent actual time looking into. Gnomingstuff (talk) 18:30, 15 August 2026 (UTC)Reply
The question is whether this will help new editors to develop good judgement. LittlePuppers (talk) 04:55, 16 August 2026 (UTC)Reply
Please just throw me in a ditch. Polygnotus (talk) 18:48, 15 August 2026 (UTC)Reply
Per WP:TALK and the header of this talk page, please avoid fact-free rants and aggressive exclamations, and try to contribute actual arguments instead. Regards, HaeB (talk) 07:01, 16 August 2026 (UTC)Reply
Fixing this error improves the comment's correctness.  Hex talk 15:46, 16 August 2026 (UTC)Reply
The above exclamation was not nearly aggressive enough. Let me rephrase for clarity: This is one of the worst features I have ever seen proposed for anything, ever. Enabling a feature like this is literally insane. –jacobolus (t) 20:10, 29 August 2026 (UTC)Reply
At the very least this specific feature will be community configurable via MediaWiki:Editcheck-config.json/Special:EditChecks once it goes live so we can turn it off. But yeah, my first reaction is the same as Polygnotus'. * Pppery * (alt) in solidarity 18:58, 15 August 2026 (UTC)Reply
Kill this with fire, and fire whoever wanted to impose this upon us. Dishraceful and going against clearly expressed community sentiment. The Wmf should not produce any tools that make content suggestions ever, this is not what they zxist for. Fram (talk) 20:08, 15 August 2026 (UTC)Reply
Kill it with fire. The only good thing about this is that it appears we have the ability to turn it off. Tazerdadog (talk) 20:54, 15 August 2026 (UTC)Reply
I fear the suggested edit feature has shown that new editors have a bad tendency to follow suggestions blindly. This isn't the fault of new editors, but poor explanations of what is being suggested and that they are only suggestions.
Looking at the example above make me think this will only make the situation worse, especially as the LLM shows that it doesn't understand policy (a common problem for ever LLM). -- LCU ActivelyDisinterested «@» °∆t° 20:59, 15 August 2026 (UTC)Reply
@ActivelyDisinterested, @Kowal2701 and Gnomingstuff, and others in this thread, you are looking at a feature in it's pre-pre-pre-alpha stage. What the ticket tells you is that the Editing team is preparing to deploy a very very early experimental version of the feature to experienced editors who have opted into enabling a suggestion mode beta and append a specific parameter to the URL (i.e. basically nobody will get this feature unless they specifically click a link and have a very specific beta preference enabled). The code is being enabled so that it can be demoed to folks, used to perform rudimentary qualitative A/B tests and gain very preliminary feedback from Wikipedians at conferences (which occurs before wider consultations with the community). This is nowhere close to being deployed anytime soon without significant bug fixes and the call-to-action in the thread, "is due to be announced any time now" is just patently false. For what it's worth, I personally haven't made my mind up about this specific feature, but I'm willing (and would strongly urge other folks) to provide the team with the ability to spend some more time atleast trying to iterate and experiment on the feature to see if some variation of it could be made useful to some Wikipedians. Sohom (talk) 22:50, 15 August 2026 (UTC)Reply
I realize that this is an experimental version of the feature, but based on the actual content that exists, this isn't the "show to editors and assume they like it" stage, it's the "internal minimum-viable-product demo" stage -- and a MVP you'd need to very carefully babysit to make sure something like The current content is a series of JSON objects that do not convey readable information to the reader. doesn't pop up onscreen. Arguably it's not even that, but the stage of "go back to square one and rethink because the premise is inherently flawed."
I think that "due to be announced any time now" is a fair interpretation of there being an August ticket called "Announce availability of 'experimental' suggestions" with the description The announcement we publish ought to equip volunteers with the info. they need to answer the following questions.... Like... it's due to be announced. That's... what the ticket... says.....
  • "Assume" is not my wording, it's directly from the ticket: In T428311 and T431376, we – staff, in collaboration with experienced volunteers across a range of Wikipedias – will have assumedly determined the initial batch of LLM-generated MoS suggestions to be reliable.
Gnomingstuff (talk) 06:46, 16 August 2026 (UTC)Reply
Let's nip this in the bud please before it becomes a fait accompli (if it hasn't already). Incredible that there's been no community consultation about this AFAICT Kowal2701 (talk, contribs) 22:35, 15 August 2026 (UTC)Reply
I'm somewhat interested in what an llm could dig up in a widespread analysis of MOS:GEO, but unfortunately the suggestions for MOS:GEO are not really about MOS:GEO but are normal typos and (misunderstood) context suggestions. This may be in pre-alpha, but it is frustrating to read "we’ve been successful in developing bespoke, one-off models that surface specific kinds of editing suggestions in a reliable way. For example, we use the Add-a-Link model to suggest relevant inline links between articles" when there have been deep flaws to the add-a-link model that have been unaddressed since its implementation. It is also known that the revert metric used is flawed, so it is disappointing to see it still being referred to. (I recently provided an example to WMF devs of a tone check edit making the article more promotional, but I don't know if that's an edge case or a more widespread issue like add-a-link has.) The "tools that show promise" user story is also quite cheeky; I can't decide where that lies on the amusement to annoying scale, it could be seen as endearing.
On the current suggestions, there is a mix of "valid and useful" and quite wrong. The valid and useful ones I saw were mostly typo suggestions. The MOS:GEO ones that went beyond that were sometimes nonsensical. The NPOV ones I checked I would avoid suggesting. The metric being used to assess readiness, in this case "The size of this vetted set of suggestions has given the team the confidence", needs to be relooked at. A consideration that seems to be lacking from all these suggestion ideas is that putting any of these into a formal structure gives them an imprimatur of authority. That's tricky to work around, but the "accelerate the speed with which we can surface meaningful signals" language suggests it isn't a strong consideration. "low-risk edit suggestions (as defined above)" is another metric that needs to be reassessed, if you're trying to touch upon NPOV you have left low-risk behind. If it is true that "it can take more than a year to produce a single type of suggestion", it does seem like far too much effort given the quality of the results. I hope the experimental team will have another think about the fundamental assumptions here and the metrics used for assessment, to help shape future development. CMD (talk) 23:42, 15 August 2026 (UTC)Reply
The idea itself is interesting, and I wouldn't be against experimenting with AI as a way to surface article quality issues (e.g., what EditCheck is doing). This seems to go much further, with the model presenting specific, actionable suggestions, although it stops short of ready-to-post edits.
Is this still experimental? Absolutely. As @Sohom Datta points out, this is about to be deployed as "experimental suggestions" open for further feedback, to a testing audience clearly distinct from its target audience. I doubt the kind of editor knowledgeable about beta features and actively seeking out to test this will be misled by the AI's suggestions.
I will concede that there is an ambiguity in the way the ticket presents the matter, which is not ideal: the purpose of these edit suggestions is to both edit more effectively with tools that show promise and contribute to making those suggestions more reliable by using and evaluating them in real editing contexts. This, while technically true, should be clarified to shift the emphasis towards the latter.
Now, what gives? Of course, no one wants the current version to be shown to newcomers, given the major flaws pointed by Gnomingstuff and others above. What we can do, however, is twofold. Now, discuss whether this feature could be developed to provide constructive help in theory, not considering its current lack of readiness. This is an open question. And later, once the feature is sufficiently mature and ready for a rollout to newcomers, discuss whether that current state is worth rolling out, whether it needs further development, or if the project should be cut short. Chaotic Enby (in solidarity · talk · contribs) 23:42, 15 August 2026 (UTC)Reply
But we've done this so many times where features seem to never get cut short after a certain point in development regardless of feedback, it's only the rare occasion when enough people kick and scream Kowal2701 (talk, contribs) 00:29, 16 August 2026 (UTC)Reply
Agree, and this is why the sunk cost fallacy has to be considered. In fact, I can see the opposite outcome from this initial discussion, namely the developers being left with the impression that all the issues the community has are with the current state of the project, and that investing more will make it worthwhile and gain community acceptance. This is in fact far from obvious, and why we should, I believe, center the discussion on the viability of the project as a whole. Chaotic Enby (in solidarity · talk · contribs) 00:40, 16 August 2026 (UTC)Reply
It's just very difficult to have a constructive discussion/approach when most are worried/anxious they're going to end up ignored and powerless Kowal2701 (talk, contribs) 01:14, 16 August 2026 (UTC)Reply
Pointing down to Peter's comment below, but basically: these specific suggestions were done in a way to try to avoid sunk costs. The development effort on Editing's side has been quite low (gerrit:1299651 + gerrit:1320209) and has mostly been about setting up a framework for a suggestion-type that can ask an API to hand us fairly arbitrary suggestions on an article and then for us to log feedback from a user about whether they seem valid. The broad idea was to make them available to people who knew how to toggle past multiple layers of "are you sure? this is a beta / experimental", and gather data about which specific instances were considered helpful and which weren't. (Screenshot below as well, but we're also not showing the content-specific suggestions you can see in the raw data file; we just show the static_description field for each one, so you just get the "this might violate MOS:GEO" level of prompting about it...) DLynch (WMF) (talk) 13:30, 16 August 2026 (UTC)Reply
I think the first wave of responses show useful feedback on what instances are helpful, what aren't, and what need extensive tuning to filter out tricky cases and focus on the most useful / confident / low-risk suggestions. Glad that NPOV is being dropped; that's extremely complex and contextual.
A good recurring issue that won't show up in spot checks of individual suggestions, is that articles overall deserve article-level checks before spending time fixing small details. @Gnomingstuff put this well above. I would be interested to see a version that allows article-level suggestions (e.g., for tags that might apply to entire articles or sections), especially for new or single-author articles.  SJ + 22:53, 18 August 2026 (UTC)Reply
I am curious what you mean by whether this feature could be developed to provide constructive help in theory. My experience as an engineer is that these sorts of discussions are not productive. You simply make things, you learn a lot along the way, some of the stuff you scrap, other stuff ends up being very useful, but not towards the goal you originally set out to achieve, and on occasion you actually end up producing what you originally set out to do. Discussions about what sort of engineering efforts would theoretically produce good products are simply a waste of time. Czarking0 (talk) 15:52, 16 August 2026 (UTC)Reply
Hey all -- I'm Marshall Miller, director of product at WMF (this project is with the teams that I work with). I'm commenting to let you know that we see this and that members of the Editing team will be able to comment with more detail, background, and clarifications this week.
But yes, let me first say that this is the very earliest stage of testing/trying/experimenting with this idea, and just for experienced editors. As we have done for all the edit check and suggestion features so far, we will only advance this feature farther if communities are supportive, if the suggestions are reliable, and if the data shows that they make a positive difference for the wiki. And all these checks and suggestions are configurable by communities at Special:EditChecks.
About why we're pursuing this: we have seen good success and community support with edit checks and edit suggestions, and the idea of suggestion mode. So far, these have run off of either simple logic ("this blob of text was pasted from ChatGPT") or small machine learning models ("this sentence uses peacock words"). What they all essentially do is point out to human editors when they are violating a wiki's policies in some way -- and they have been shown to reduce revert rates and make newcomers more successful. But there are a lot of wiki policies that could be useful to point out to people. LLMs are a new technology, and they can and do make mistakes. It takes careful testing and tuning and evaluation to get them to perform reliably and may not even always work out (and it may not in this case either). But they may make it possible to produce more of these useful checks that help newcomers make better edits, and may help experienced editors notice things that need fixing. We will 100% need the input of all of you to help us figure out together whether we're on to something or not.
Our approach to all this is that editing decisions should be made by humans (except for the very simple kinds done by things like ClueBot, etc). And that features like these try to help humans notice/find places where they could apply their judgment.
Okay, anyway -- more to come from team members who are deeper in the details. MMiller (WMF) (talk) 04:03, 16 August 2026 (UTC)Reply
WMF A/B tests are notoriously unreliable, and invariably interpreted in the most positive light possible. See e.g the image viewer disaster, or the initial claims about edit check where the posted positive results turned out to be false. Why should we trust whatever results will be posted this time? More importantly, why is such a tool created when the WMF should know by now the massive pushback they would get against AI content suggestions? Aren´t there enough other improvements requested (often for many years already?). Fram (talk) 06:51, 16 August 2026 (UTC)Reply
So far EditCheck has produced amazing results (e.g. vastly increased the share of newcomer edits that contain citations) – and the Editing team listened a lot to community members while developing new checks/suggestions. But you don’t have to trust anyone given that all suggestions and edit checks can be enabled and disabled by local admins. Johannnes89 (talk) 16:56, 16 August 2026 (UTC)Reply
Yeah, I think I meant referencecheck (or whatever it is called), not editchecks (which I haven't checked), should have been more careful in what I said. They claimed a serious number f added references, but it turned out that a lot of edits were tagged as "reference added" when this wasn't true, and a lot of other "references" were completely invalid but counted as a success anyway. I posted this with clear examples, but the WMF ignored this completely. But that's about reference adder AB tests, not edit check, so again, I should have checked before posting. Fram (talk) 09:41, 17 August 2026 (UTC)Reply
From glancing at the list linked above (totaling 349 suggestions - 22 labeled MOS:GEO, 116 labeled NPOV, and 211 labeled "simplify language"), here are my thoughts:
MOS:GEO - most seem fine at a glance. Lots of suggestions to fix diacritics and spelling. Several which are well outside the purview of MOS:GEO. Seems lacking in nuance in some edge cases (but saying more confidently would require some fact-checking).
NPOV - lots of issues. It is too timid to say anything forceful, even when it's warrented and supported in RS, and wants to add qualifiers (e.g. "reportedly") for simple statements of fact, such as "improved quality of air" or "top of the chart". It seems generally opposed to any words which are not incrediby boring, and even some which are: perilous, successful, unreliable, transparent, vocal critic. It is often very unclear in what it is referring to ("the evaluative phrase", "the promotional claim", "subjective description", "the evaluative language"). Multiple suggestions are to remove language which is not there. It also has a terrible time recognizing attribution, and suggests several times (probably a dozen+) that it be added when it's already there. I've skimmed through maybe half of these, and a majority have issues.
Simplify language - meh. Some are fine. It seems to want incrediby short sentences. Ironically, I also disagree with the one I see where it suggests combining sentences. Most suggestions are pretty vauge. One suggestion is "make this neutral". Also says "American English is preferred on Wikipedia" on a British biography.
Summary of my views: GEO is mostly decent but may lack nuance, NPOV isn't really useful because it lacks the understanding to know when a strong viewpoint is neutral and generally dislikes big words, and simplify language is overzealous in suggesting short sentences. LittlePuppers (talk) 06:16, 16 August 2026 (UTC)Reply
Okay, I was looking at a different file from Gnomingstuff and one which is at least a few weeks old. Take that how you will. LittlePuppers (talk) 06:19, 16 August 2026 (UTC)Reply
Glancing through what is (I think) the latest (and much longer) version, there may be some improvement, but most of my thoughts still apply. LittlePuppers (talk) 06:29, 16 August 2026 (UTC)Reply
The MOS:GEO ones are often not fine. For a start, being outside the purview of MOS:GEO suggests some underlying flaw in the model. "Update the country name to conform with Wikipedia’s geographical naming conventions", I have no idea what that is meant to refer to. "The sentence is amended to specify that the Grand Canal is in Venice, providing clearer geographic information" lacks understanding that the Venice location was established in the prior sentence. "The parenthetical abbreviation after “Guantanamo Bay detention camp” is incorrect and should be removed" is simply wrong, although it is perhaps an unnecessary abbreviation a reader may also not understand. "The place name "South Island of New Zealand" is not formatted according to MOS:GEO; it should use commas between the island and the country" is again just wrong. "The original text mentions "the Atlantic" without specifying that it refers to the Atlantic Ocean, which may cause confusion", not sure what to say about that one, Atlantic Ocean is even written out explicitly earlier on the page. CMD (talk) 06:35, 16 August 2026 (UTC)Reply
Yeah, the set I was looking at initially had a very limited list for GEO. A lot of what you mention reflects broader issues with all the categories as well. LittlePuppers (talk) 06:56, 16 August 2026 (UTC)Reply
Most of these are known issues with LLM-suggested edits, or at least the kind of thing that has certainly been possible to know about for at least a year:
  • The "promotional claim"/"evaluative language" stuff is a 2025-era LLM tic. Very specific verbiage, especially the "evaluative" part, that shows up over and over again in AI edit suggestions and basically nowhere else. Here's a bunch of examples.
  • The "original text mentions the Atlantic" suggestion is another consequence of the isolated-sentences approach; the LLM is responding to the sentence "According to Herodotus they dwelt geographically along the sea south of Libya on the Atlantic," and so it doesn't have the context of first reference/subsequent reference.
Gnomingstuff (talk) 06:56, 16 August 2026 (UTC)Reply
Further context or not, saying "the Atlantic" is not going to cause confusion. CMD (talk) 07:06, 16 August 2026 (UTC)Reply
"American English is preferred on Wikipedia", great so it's not just wrong but will make a bad situation worse. -- LCU ActivelyDisinterested «@» °∆t° 09:07, 16 August 2026 (UTC)Reply
The issue with language suggestions in regard to NPOV is that LLM are not neutral, and do not give neutral suggestions. So using them to make these kind of suggestions is a way of creating a fake consensus. -- LCU ActivelyDisinterested «@» °∆t° 09:11, 16 August 2026 (UTC)Reply
Thanks Gnomingstuff for this thorough demonstration of why LLMs are systems for producing text-like slop. No matter how many attempts are made to patch these behaviors, they will keep happening because no comprehension or intelligence is involved, and never will be. Just guess after guess, a fountain of hot slop staining our precious reputation as one of the few uncontaminated places online.
This shameful and embarrassing effort needs to be canceled immediately and the donation money wasted on it so far written off. I'm not even going to start getting into the unethical nature of using LLMs in the first place, which should have been sufficient on its own to rule out even considering something like this.  Hex talk 15:58, 16 August 2026 (UTC)Reply
The sheer quantity of bullshit from the WMF is exhausting at this point. Cremastra (talk · contribs) 04:41, 17 August 2026 (UTC)Reply
This is yet another reason to not trust the WMF and it shows how it's impossible to assume good faith on their part. The WMF at this point is an active threat to the very existence of Wikipedia. Ita140188 (talk) 07:58, 17 August 2026 (UTC)Reply
Dealing with WMF is like living through Groundhog Day. They have too many employees so they bureaucratically create "jobs" building crap that nobody asked to solve "problems" that don't really exist, creating a bigger set of unforseen consequences (because WMF is composed of many software engineers and few Wikipedians and is always and forever tone-deaf to community desires). The volunteers who make the project run are all "power users" to them... Well, here's what the "power users" are saying, "tech bros"...... NO AI ON WIKIPEDIA. Didja get that? Carrite (talk) 14:57, 17 August 2026 (UTC)Reply
Wikipedia has a reputation as one of the last bastions of information, in an era of hallucinated, enshittified LLM-generated slop. Any embrace of AI-powered anything on the platform constitutes a plan to throw that into the bin. ser! (chat to me - see my edits) 11:51, 18 August 2026 (UTC)Reply

Taking a step back: could AI suggestions be beneficial?

[edit]

As pointed out above, the current development stage is way too early for a broad rollout. This should have been better clarified, both to reassure the community about the experiment, and to provide clearer development goals.

However, taking a step back, a discussion can still be held regarding the potential of this whole endeavor. Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way, assuming the current pitfalls could be smoothed out in development? More concretely, are these expectations realistic, and is it worth investing further resources in this project?

These are questions I don't, personally, hold the answers to. However, we should be discussing them together, alongside members of the Editing team involved in its development (courtesy ping to @Quiddity (WMF)), if we want them to work in sync with community sentiment, and avoid investing resources in dead ends. Chaotic Enby (in solidarity · talk · contribs) 23:57, 15 August 2026 (UTC)Reply

Would the community, in theory, agree to an AI model surfacing suggestions to newcomers. I think a better way to explore this tool would be to make it available to established editors first. People who have the experience and policy knowledge to be able to properly evaluate the suggestions. Maybe the people will say "The suggestions were all spot-on and incredibly valuable". Maybe they will say "Nothing this thing suggested made any sense at all, it's total garbage". More likely, somewhere in between. But let's do the experiment rather than pre-judging it. RoySmith (talk) 00:08, 16 August 2026 (UTC)Reply
I think a better way to explore this tool would be to make it available to established editors first. While it might not have been clear at first, this is, in fact, exactly what the experiment is planning to do. The tool is still in development, and we can't say, in advance, how it will end up in terms of quality. For now, I'm just trying to figure out the proportion of editors who either will find it a non-starter in principle (regardless of the suggestion quality) or are opposed to investing further resources in its development for any other reason. Chaotic Enby (in solidarity · talk · contribs) 00:20, 16 August 2026 (UTC)Reply
Mark me down as "opposed to investing further resources in its development for any other reason". Wikipedia has a large amount of technical debt that would be easy for a WMF developer to fix, but instead they are doing this?
As one example, Arbcom is a very important function, and many people who participate get stressed over whether they are under the word count limit. But the tool that puts a banner at the top of your comment fails to accurately count your words! Worse, you can't invoke it while composing -- you have to post and hope that you didn't go over. And if an arb replies to you inline, that increases your count! This is the sort of thing that a WMF developer could fix in an afternoon. It's important, but making an accurate arbcom word counter will never hit the top of any survey of things many editors want to see fixed.
Another example: what happens if when I sign this comment I accidentally hit the "~" key three times instead of four? How about 5, 6 or 7? This is a typo that happens again and again. How hard would it be for a WMF developer to make it so that you get a "are you sure" message before accepting a signature that is almost always a typo?
There are hundreds and hundreds of these easy to fix things that depend on old scripts written by volunteers and all too often no longer maintained.
I think the WMF should spend a significant amount of developer effort -- 75% or 80% -- fixing these small, non-sexy quality of life issues and only then devote the other 20% - 25% to fun things like AI suggestions. --Guy Macon (talk) 01:02, 16 August 2026 (UTC)Reply
Or, if you don't like the above 4-tilde signature, --Guy Macon (talk) (3 tildes),  --01:13, 16 August 2026 (UTC) (5 tildes), or --01:13, 16 August 2026 (UTC)Guy Macon (talk) (7 tildes).Reply
Guy, on that last one, you don't need to add a signature at all anymore. Discussiontools will do it for you. I haven't added one to the end of this post, for example. In solidarity, asilvering (talk) 01:04, 16 August 2026 (UTC)Reply
Will I (Guy Macon) get an error message for this unsigned post or do I have to make a preferences change to get the autosign goodness? Show preview says it will post the unsigned comment with no error message.
Discussiontools is a nice tool, but it is no substitute for software baked into Wikipedia that checks for common errors and throws up an "Are you sure?" message. I could give you a hundred examples of places where the WMF is depending on unpaid volunteers to maintain basic functions that keep the Encyclopedia running smoothly while focusing on the exciting new stuff --01:30, 16 August 2026 (UTC) Guy Macon (talk)Reply
Guy would need to be using DiscussionTools for that, but unfortunately they are using classic full page source editing where none of those helpful things will occur. DLynch (talk) 13:13, 16 August 2026 (UTC)Reply
FYI I made Module:Word count, I can refine this further if you think it would be useful. It does not 100% solve the issue you mention Czarking0 (talk) 15:57, 16 August 2026 (UTC)Reply
Although I agree with a lot of this sentiment, I think an organizational strategy which places say 20% of the engineering resources on long term tech rather than present issues is reasonable. As long as the development of this feature counts under that I do not see the problem. Czarking0 (talk) 16:00, 16 August 2026 (UTC)Reply
+1, that seems to be what is happening here (see the comment below about 'avoiding sunk costs' and getting feedback early and often from established editors). I can see individual categories of suggestion being useful to experienced editors, focus on making something that works for them before considering anything that might be visible to newcomers. (They already have Special:Homepage)  SJ + 23:14, 18 August 2026 (UTC)Reply
I agree with Fram above in that this feels like a way (though a much more subtle way than Simple Summaries) for the WMF to influence editorial decisions which, with the exceptions of legal reasons, they just should not be a part of. JCW555 (talk)01:55, 16 August 2026 (UTC)Reply
There isn't really any plans for the WMF to get involved in editorial descision (and I say that as somebody who has through m:PTAC reviewed the Annual Plan). The plan for Edit Suggestions is purely meant as a assistive tool to help editors and kinda comes from the idea of being able to surface gadget/userscript suggestions to everyone without having to have coding knowledge. Sohom (talk) 02:57, 16 August 2026 (UTC)Reply
But the fact that the AI is making these suggestions at all in the first place is the WMF having a subtle hand on editorial decisions in my mind. The examples Gnomingstuff lists above, like the FreedomWorks example, is an example of the AI making an editorial judgement that it should not be doing. Some of the others are more subtle like the DeAndrea G. Benjamin example, but they're still editorial decisions that the WMF shouldn't be engaging in period. JCW555 (talk)03:23, 16 August 2026 (UTC)Reply
They are using generic models without fine tuning for these suggestions, so out of all the parties that could be said to have made editorial decisions in this scenario I would say OpenAI and Google would rank above the WMF, and I really wouldn't consider either of those companies to have made any editorial decisions... the problem IMO is potentially encouraging uh... not really making any editorial decisions, and vibing through things without considering the context of, e.g. the contents of the sources as we are supposed to. Alpha3031 (tc) 08:28, 16 August 2026 (UTC)Reply
It is not worth creating AI prose suggestions for newcomers (especially if it takes a year for each model!). Putting aside quality questions, editors here are expected to be competent and able to contribute in English. While there are different ways to be confident, we expect the ability to read and write in English, and thus to some extent to have the tools to be able to copyedit themselves. We also expect editors to be able to read and comprehend our guidelines and policies (pre-emptive clarification, read, not memorise). The best way for us to be able to understand competence in these areas, and thus to be able to assess and offer advice if needed, is to see their edits. Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors, and harmful to the new editors in giving them a misleading picture of how things work, and more harmful when the llm doesn't understand our policies and practices (as this one does not). Apologies to Sohom, but "There isn't really any plans for the WMF to get involved in editorial descision" just isn't true if this sort of system is being set up. This sort of suggestion task is the WMF making editorial decisions, even if they're making it through an llm. There are many reasons time would be better spent anywhere else. (I distinguish "prose suggestions" from say the add-a-link task, as that at least teaches a technical competence that we do not expect new editors to have. Such teaching does seem useful, although as mentioned above it would be nice if it was developed further.) CMD (talk) 03:45, 16 August 2026 (UTC)Reply
Seeing instead a whole slew of new editors making the same llm-prompted changes is harmful to the community in being able to understand and accommodate new editors I just want to emphasize this -- the English Wikipedia community is not good at welcoming newbies who make mistakes. That is a problem.
The English Wikipedia community is actively hostile to editors making mistakes with large language models.
If a newbie puts a poor LLM-based reword into an article, an experienced editor is almost certainly going to WP:BITE them off. And a newbie, operating in a system they're unfamiliar with, with a human-sounding voice telling them "This copyedit is right", is never going to have the knowledge and is almost certainly not going to have the confidence to challenge the AI when it presents them with a bad suggestion. That is going to set them up for failure when a grumpy human editor, burnt out from dealing with LLM edits, callously reverts them.
I think using machine learning to help with encyclopedia maintenance is a wonderful thing! And I think large language models are really cool -- but they require a high degree of skill to use correctly in the manner that looks like it's being explored here. And, again, the community is so burnt out from dealing with poor quality LLM content that any editor who uses this tool and (inevitably) makes a mistake will be attacked by the community. Does the team working on this understand that, @MMiller (WMF)? GreenLipstickLesbian💌🧸 05:10, 16 August 2026 (UTC)Reply
Seconding that this could be helpful for all sorts of maintenance but should be aimed at experienced editors while working out kinks.
I would personally like to see rubrics for highlighting potential vandalism and LLM edits. And I want much less text in my sidebar and zero suggested text: just a few words indicating the kind of issue to look for / the kind of style guidelines to check.
And I'd like to see explicit self-evals of each rubric for suitability to the task and for false positive/negative rates, which could also be compiled and developed by community maintainers.  SJ + 23:52, 18 August 2026 (UTC)Reply
@GreenLipstickLesbian -- I think that the most important thing that would prevent against the situation you're describing is that the suggestions wouldn't propose text for the newbie to accept/reject. It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?" -- it would not give them a re-written sentence to add. I know that in that situation, the AI may be wrong about whether the sentence needs to be rewritten, and the sentence may be perfectly fine -- but the newbie might be like, "Hmmm, well I guess I'll reword it?" and they may make a pointless edit, or may make the sentence worse.
Suggestion tasks like "add a link" and "add an image" actually do give newbies specific edits to accept or reject, and we see them generally apply good judgment and be constructive. Yes, many of them mess up and get reverted, but that might have happened to them anyway if they were left to their own devices. So these tasks cause a bunch of things to happen at once, and the question is whether it all adds up to a net positive or net negative, you know?
What do you think? MMiller (WMF) (talk) 21:13, 19 August 2026 (UTC)Reply
@MMiller (WMF) Thanks for the response! Yes, I think these sound like reasonable precautions. I do really want to emphasize the point about making it clear to experienced editors what the newbies actually are seeing is very important.
Having had the beta experimental version of suggested edits enabled for a few days, I definitely see the vision. And I see how it could be very useful. (My response to many of the "consider adding a citation" suggestions has, admittedly, been a)"lol no, I'm removing the unsourced text for other PAG issues", b)"this is cited, it just needs an inline citation", c) "... yes, i see how citogenesis is made", or, d)"... we need an easily accessible essay on 'how to source a statement' that we can link to the newbies here" because wow, i'm having trouble". ) But I also see the biting that happens when newbies do make pointless edits/less than ideal edits ( has some which took more than a few years to rectify), and so I am worried about adding any "they're just thoughtlessly adding machine suggestions to articles"-esque ammunition, even it's it's not strictly true. GreenLipstickLesbian💌🧸 21:33, 19 August 2026 (UTC)Reply
This is why I suggested some basic FAQ type popups upon first edit, including an agreement to not use LLMs, to prevent true good faith mistakes and remove plausible deniability for the rest. ChompyTheGogoat (talk) 17:37, 26 August 2026 (UTC)Reply
RE: especially if it takes a year for each model! to be fair to the teams involved the stated idea behind this specific experiment seems to be wanting to try something that won't take a year per task on what the editing and ML teams think are relatively low risk tasks... I'm just not sure that the community and the teams involved have a sufficiently compatible idea of what tasks are low risk. I think it would be possible to develop something acceptable to the community, but unless very sure about it, it may be best to get a vibe check from the community before thinking something is "low risk", and treating things as "high risk" otherwise. Alpha3031 (tc) 08:36, 16 August 2026 (UTC)Reply
Would the community, in theory, agree to an AI model surfacing suggestions to newcomers in such a way I won't, and if that ever happens I'm gone. Models are bias black boxes, and even if suggestions were 100% accurate this would still be an issue. fifteen thousand two hundred twenty four (talk) 04:37, 16 August 2026 (UTC)Reply
If AI had far more quality control, I wouldn’t be opposed to this. Except it doesn’t yet. LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician.
This could probably work in theory. But it would assume an AI with a reasonable knowledge of the given article subject, some common sense, and a degree of fluency in English (by which I mean not doing things like pyusician/musician), and, given the above examples provided by Gnomingstuff, this is definitively not that.
I oppose AI-based features being integrated into Wikipedia software, and will continue to do so unless a day comes when LLMs equal or surpass the common sense of a human. Perhaps in a few years LLMs will be reliable enough for this to work smoothly and without issue. With the current state of AI, I don’t think we’re there yet. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 04:59, 16 August 2026 (UTC)Reply
The AI could work 100% of the time and I'd still oppose it because these edit suggestions are being directed by an AI that's controlled by the WMF, which outside of legal purposes, should never have any editorial control on Wikipedia in the first place. JCW555 (talk)05:17, 16 August 2026 (UTC)Reply
That’s a fair point; I hadn’t thought of that. (Yet another reason why this is a terrible idea.) Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 05:39, 16 August 2026 (UTC)Reply
The solution for this is making prompts publicly available and allowing each community to tweak them. Alaexis¿question? 06:36, 16 August 2026 (UTC)Reply
LLMs hallucinate, and those hallucinations are not something you want informing newcomers who are near-clueless as to how Wikipedia works, let alone people learning English who won’t be able to tell when an AI-based edit suggestion system instructs them to insert a grammatical error or something similar, like the above pyusician —> musician. - that's a rather odd example to get hung up on, given that automated spell checkers and their admittedly sometimes absurd correction suggestions have been around for decades (long before the term "hallucinations" came to be used in this context). Many or even most professional writers - journalists, book authors, scholars - use them routinely, and have learned to live with the occasional fail of this type (pyusician —> musician).
Indeed, in stark contrast to your logic here, WP:SPELLCHECK has long described them as potentially useful for Wikipedia editors, too, as long as they don't blindly rely on such tools:

Spellchecking software and online tools can be helpful when copyediting Wikipedia articles. [...]

  • No spellchecker is completely accurate. You must check the output of any tool you use. [...]
  • You are responsible for all spelling and grammar changes you make, even if the changes are suggested by an error-checking tool, such as Grammarly or ChatGPT.
Maybe it is time to conceive of Wikipedia editors a bit more as adults who are generally capable enough of using such tools even though their suggestions are not 100% error-free (very few things are).
That said, I do generally agree with your point that AI needs quality control, and folks should definitely ask if WMF has done enough here yet (I'm not sure it has). It's just that - as the spell checker example shows - it's not realistic to demand 100.000% accuracy, or to point to isolated failure cases without assessing how frequent they are. (To be fair, User:Gnomingstuff did already make an informal heuristical effort at the latter with regard to their list above: it took me only about ~1-2 hours to find the above issues, and that was only a spot check, i.e. there is informal evidence that these are not very rare at this point. But ultimately I'd be interested in more concrete assessments of how likely editors using this tool will be to encounter each of those failure cases in practice.)
Regards, HaeB (talk) 06:54, 16 August 2026 (UTC)Reply
Most adults have a lot more experience with spelling than they do with Wikipedia's guidelines. LittlePuppers (talk) 06:58, 16 August 2026 (UTC)Reply
Sure, but that doesn't mean that we keep the edit button away from them.
See Wikipedia:Competence is required, which is perhaps a better known version of this principle that we generally expect editors to be competent adults (metaphorically, with apologies to all the very smart and capable teenage editors among us) that do not require special restrictions and safeguards to protect them against their own mistakes.
Regards, HaeB (talk) 07:30, 16 August 2026 (UTC)Reply
The thing I’m concerned about is everyone knows that spellcheck is obviously not infallible, as the Internet often likes to humorously point out. But to a newcomer, anything that comes from ‘Wikipedia’ seems naturally correct and reliable.
Certainly not everyone would fall into this trap, since, as you point out, it isn’t like all newcomers are to be treated as children who don’t know what they’re doing. But it’s likely that some would; perhaps enough that the consequences of Wikipedia’s reputation for factual accuracy is something we should take into account on this. Given all of that, I don’t think we can draw a one-to-one comparison between run-of-the-mill spellcheck and an edit suggestion feature built into Wikipedia itself.
That’s why I argue that we should hold ourselves to a higher standard; after I’ve given your points some thought, however, I think you are correct that my own request for equal [to] or surpass[ing] human-level reliability is asking a bit much of a literal artificial intelligence. Simultaneously, I think spellcheck-level absurdity is far too low a bar for this feature.
Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 09:11, 16 August 2026 (UTC)Reply
Sounds like something that could be addressed with a user level disclaimer about LLMs before the feature is enabled on their account. Czarking0 (talk) 16:02, 16 August 2026 (UTC)Reply
That would definitely fix that issue (I’m surprised I didn’t think of that, to be honest). Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 00:11, 17 August 2026 (UTC)Reply
I agree with @RoySmith that some suggestions are more suitable for experienced editors. I believe that checking citations is a good use case with an LLM flagging potentially problematic citations and editors verifying them manually (we have a proof-of-concept that I've been maintaining but as long as it's a userscript it won't make a dent in the sourcing problem). This is a task that is not too complex but requires familiarity with WP:RS and some experience. Alaexis¿question? 06:35, 16 August 2026 (UTC)Reply
Based on the suggestions here, my stance is basically the same as what WP:LLM already says: typos and punctuation suggestions seem unnecessary but mostly fine (except when they're not, see the "physician" example), anything beyond that is unlikely to be fine, and suggestions related to NPOV are light years away from fine. 100% accuracy isn't possible but, realistically speaking, people are going to rubber-stamp whatever suggestions they receive.
As far as the spot-check goes, the timeframe is not really all that scientific since I actually saw this last night (not even while doing AI stuff, I was doing Commons new image patrolling and saw one of the screenshots from it), slept on it, and did the full writeup the next day. I also have seen tens of thousands of AI edit suggestions, as well as thousands of snippets of parsed Simple Summary text, so I already knew the places where this was likely to run into issues. This is also why I think the "experienced editor"/"new editor" dichotomy is not helpful, in general. Experienced editors are not necessarily experienced in copyediting and/or the specific problems that crop up with AI, and newcomers are not fragile baby birds who have never edited a piece of writing in their lives. Gnomingstuff (talk) 07:10, 16 August 2026 (UTC)Reply
Automated suggestions could be helpful in very limited circumstances. I can imagine an experienced editor who has chosen to opt in to them assessing each proposal carefully, skilfully rejecting those with the sorts of problem listed above, and improving Wikipedia by acting only on the good ideas. I'm sceptical that AI is the best way to produce such input but I'm willing to be proven wrong. Methods such as this are totally unsuitable for new editors, many of whom will blindly obey The System, make bad edits in good faith, and get reprimanded or blocked. There is also the unoriginal but valid argument that editors would prefer the WMF to devote its resources to other activities, even if we don't always agree exactly on which ones. We're seeing a lot of these "Here's a new toy no one asked for that we've secretly blown this year's development budget on" announcements. I do honestly try to hope for the best, but I'm afraid my initial reaction has become "Oh no, what have they broken now?" Certes (talk) 09:57, 16 August 2026 (UTC)Reply

Wouldn't it make more sense to focus on basic copyediting suggestions? I just went to a random article, Mircea II of Wallachia, and requested some copyedits from chatgpt.

Most of these seem like pretty decent suggestions, don't dip into npov/tone danger zones, and are simple enough for new editors to review and handle. ScottishFinnishRadish (talk) 00:14, 20 August 2026 (UTC)Reply

Also, that might be the first time my first random article click wasn't about a sports ball player. ScottishFinnishRadish (talk) 00:15, 20 August 2026 (UTC)Reply
the nice thing about having an insite spell/grammar checker would be less people would use Grammarly (we could tell people using Grammarly to use the insite one instead, similar to how PasteCheck works). I've found asking a chatbot for corrections gives decent results, but obv you can't defer to them (esp. on content) Kowal2701 (talk, contribs) 19:59, 20 August 2026 (UTC)Reply
Why is it a problem that people use Grammarly? RoySmith (talk) 20:02, 20 August 2026 (UTC)Reply
it goes beyond its scope as a spell/grammar checker and rewrites content, and a lot of people using it don't know it's LLM-powered. It's the worst because people will put time and effort into writing, then put it through Grammarly which turns it into slop. We see it at AINB a fair bit Kowal2701 (talk, contribs) 20:10, 20 August 2026 (UTC)Reply

Context from Editing Team

[edit]

Hi y'all – I'm Peter Pelberg, product manager of the Editing Team. Together with the Machine Learning Team, we've been working on the experimental model-generated edit suggestions. You have raised a range of valid concerns/questions that warrant responses. You can expect those when I'm back online in earnest next week.

In the meantime, I'd like to clarify some aspects of the work and the thinking that's informing it.

First and most importantly: there are no plans right now to deploy LLM-generated edit suggestions as default-on to anyone. The closest thing to a deployment that we've talked about is a potential A/B experiment that would not move forward until we, volunteers and staff, have deemed these suggestions reliable and promising. This commitment is in Phabricator by way of the experiment (T431377) needing T428311 to happen first. For context, T428311 states, "Learn whether experienced volunteers at en.wiki think the experimental suggestions are sufficiently reliable to be shown to newcomers so that we can decide: Will we invest the effort to scale this initial set of LLM-generated MoS suggestions across languages via T431376?"

Now, with regard to what this initial set of 30,000 LLM-generated suggestions is and is not…

  1. Who has access to these experimental suggestions? In this initial, experimental phase, these suggestions will only be available to people who A) enable the Suggestion Mode beta feature and B) install this user script. Note: We will soon introduce a setting within Special:Preferences so people interested in trying out experimental suggestions don't need to install a user script to access them.
  2. What do these experimental suggestions do? This initial batch contains experimental suggestions for three types of improvements/issues (listed below). Each suggestion will highlight the span of text it is relevant to, offer a generic description of the issue, and crucially, ask the experienced volunteers encountering it to indicate whether you think the suggestion itself is valid/useful or not. And if not, offer an explanation as to why. Said another way: These suggestions do not offer fixes or ask experienced editors to make any.
  3. What are "experimental" suggestions? "Experimental" suggestions are a set of – as I think @Sohom Datta put it well – pre-alpha suggestions. The purpose of them is to evaluate their reliability. Currently, there are a total of 6 experimental suggestions available at en.wiki to people who opt-in to seeing them. 2 of these 6 suggestions are powered by machine learning models; the other 4 are based on deterministic heuristics. You can see the full list here: Special:EditChecks#Experimental checks.
  4. What types of suggestions is this LLM generating? This initial batch contains suggestions for simplifying language, rewriting language in a neutral point of view, and adjusting place-based names to follow Wikipedia's geographical guidelines. We are by no means committed to this set of suggestions or to using LLMs in general. We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk.

Last thing for now: we are aware of the sunk cost fallacy. Thank you for raising it, @Kowal2701. In fact, as I hope the above demonstrates, the whole point of developing experimental suggestions, and making them available in production to experienced volunteers who have explicitly opted into them, is as @Chaotic Enby described: to assess whether they're reliable. Ones that aren't, we will abandon. The ones that are, we'll work together to figure out how and to whom to make them available.PPelberg (WMF) (talk) 05:48, 16 August 2026 (UTC)Reply

as […] “1.” and “3.” demonstrate. This is perhaps not the most important point, but it is bugging me: were the numbers meant to all be “1.”? Comment struck due to this being fixed; apologies. Cheers, 𝔰𝔥𝔞𝔡𝔢𝔰𝔱𝔞𝔯 (𝔱𝔞𝔩𝔨) (any/all) In solidarity. 06:08, 16 August 2026 (UTC)Reply
Thanks for responding now and letting us know you don't have time to fully engage immediately, hopefully you will have time to go through all the comments above later in the week as you say. As the process of contingent experiments has been brought up again, perhaps it would help to answer the experimental question plainly, because it can be very easily answered without a complicated multi-month experiment. The suggestions are not reliable, and definitely not reliable enough to show to newcomers (not that this is the only metric that could be considered). To look further, re We chose this initial set because we found them to be reasonably accurate and assumed they would be relatively low-risk, the linked section lacks an assessment of accuracy or risk. Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers. There is a mismatch in understanding somewhere. CMD (talk) 06:18, 16 August 2026 (UTC)Reply
Part of the issue here is that a simple glance is all that is needed to identify issues with many the suggestions, so it feels like even that is not being done before these are farmed off to volunteers.
Honestly can't put it any better than that. I do appreciate the response and understand it is the weekend. Gnomingstuff (talk) 07:18, 16 August 2026 (UTC)Reply
Tbh this is a systemic issue to do with WMF governance rather than a reflection on anyone here, where developers and idea labs seem to typically be distant from the projects and communities, and everyone at both ends is left straining to compensate for this. Phab being largely open to the public goes a long way, but I wonder if there could be something like a WP:VPI on meta for developing/screening ideas (obv staff's medium is meetings and private convos etc. and idrk what currently happens, but something like "John and I were talking yesterday about this, X is a potential solution" would be good to get some feedback at the earliest stage). I've previously suggested that staff should be encouraged to do a little bit of editing each week at any wiki they choose (and reduce hours/week a bit for this while keeping salaries the same), but it risks the WMF's status as a 501(c)(3) organization (I was told) Kowal2701 (talk, contribs) 07:57, 16 August 2026 (UTC)Reply
I don't really know that it has anything to do with governance in this case. The issue would be the same regardless of whether the development process was public or semi-public or completely closed off; it's a classic, perennial issue of workplaces. The issue being -- and I am really trying to be as polite as possible here -- that QA is either not being done, or not being done correctly.
From what I understand, the suggestions were assessed for quality via AI, and then someone seems to have reviewed 15 suggestions manually. Unless there's more internal discussion (there probably is, this stuff is nearly impossible to follow), that seems to be the whole QA process.
Meanwhile, I actually have worked in QA. For a task like this, we would be given a spreadsheet like the ones here and would review every single line individually. We caught a lot of errors this way, the final product was better as a result, and it only took a day or two of work at most. I do not feel this is an unreasonable amount of work to be expected from a professional organization. Gnomingstuff (talk) 17:52, 16 August 2026 (UTC)Reply
A stage of basic competitive fault-finding, annotating a large sample by hand to flag and classify problems, goes a long way. Your immediate reflex to do this as the first response above demonstrated this nicely.
If an originating team works in public to catch and patch issues in that way, before others run into them, with high standards for accuracy, that would set discussions like this off on a better foot.  SJ + 00:23, 19 August 2026 (UTC)Reply
Thank you. Sorry, I didn't know who was in the team working on it, I trust all these people. Kowal2701 (talk, contribs) 06:51, 16 August 2026 (UTC)Reply
I oppose any A/B tests of this feature. It is antithetical to what Wikipedia should be, and is not something the Wmf should attempt. Suggestions sent by some central, biased, singleminded entity diminishes the wide variety of input, style, viewpoint, ... that makes the richness of the community. Wikipedia ia a beacon of human diversity, an LLM is the opposite of this. Fram (talk) 06:58, 16 August 2026 (UTC)Reply
In contrast to such hyperbolical rhetorical flourishes, the community has already been widely using AI suggestions sent by some central, biased, singleminded entity controlled by WMF for over a decade, in form of ORES (or now the "revert risk" models). These AI suggestions (on whether to revert another editor's changes as vandalism) are made available to any editor at Special:RecentChanges even though they might be seen as more consequential on average than those of the new tool under development here, and have a fairly high error (false positive) rate. I have personally implemented many thousands (probably tens of thousands) of those AI suggestions over the years, and survived skipping many more mistaken suggestions.
I don't want to entirely dismiss concerns about control by the Foundation though. In case of these vandalism detection AI models, the more recent work on them seems to have been dominated by decisions and aims of the WMF Research department that may not entirely align with the community here (for example, they seem to have foregone possibly substantial quality improvements for English Wikipedia in favor of language equity, and implemented their own conceptions of "fairness" with regard to IP editors).
The main employee behind the success and broad community acceptance of the original ORES left WMF years ago, and his parting recommendations for "Community-centered Evaluation of AI Models on Wikipedia" do by and large not seem to have been taken up by WMF.
Regards, HaeB (talk) 08:36, 16 August 2026 (UTC)Reply
As I said earlier, keeping the prompts (and the pipeline in general) publicly available and enabling each community to tweak them would go a long way in assuaging these concerns. Some competence is required to maintain these tools, but there is no reason not to make prompts and benchmarks transparent. Alaexis¿question? 10:50, 16 August 2026 (UTC)Reply
The system prompts used for the model appear to be published on gitlab here according to the mw:VisualEditor/Suggestion Mode/Model-generated editing suggestions#Research findings. Not sure which size Gemma gemma4:latest is, maybe the E4B? I would be interested to know if other models were evaluated, for example Nemotron 3 Super is a similarly sized model to gpt-oss:120b, but a bit newer. Mistral is dense so might be a bit slow to run. In any case, there are many newer models in the same weight class, surely it wouldn't be that hard to generate say ~1000 from each of them to see if anything is clearly better, given the only thing that has been modified is the system prompt? Alpha3031 (tc) 12:22, 16 August 2026 (UTC)Reply
Thanks. Going off of that, it appears that no reference information is given, which is not surprising looking at the NPOV suggestions, and that no actual Wikipedia policies/guidelines are included in the prompt. LittlePuppers (talk) 15:24, 16 August 2026 (UTC)Reply
A successful model would likely go beyond a prompt alone and be equipped, at the very least, with retrieval-augmented generation capabilities in order to access the text of the policies/guidelines themselves, and, in the case of NPOV, ground its answer in available reliable sources. I don't think that would be enough for NPOV specifically (there is still too much of a black-box effect in the model's weights, that won't be fully offset by prompting or sourcing), but this is to say that the current approach is far from optimal. Chaotic Enby (in solidarity · talk · contribs) 15:46, 16 August 2026 (UTC)Reply
The particular challenge is, I think, that there seems to be a trend to use increasingly sophisticated models to push increasingly challenging tasks increasingly to newer editors. LittlePuppers (talk) 15:33, 16 August 2026 (UTC)Reply
Regarding point 4., community feedback makes it pretty clear that anything regarding NPOV or tone will not be seen as low-risk, and might be interpreted as an attempt by the WMF to influence editorial decisions. I believe it would help the prospects of the experiment to commit to follow the emerging consensus, which is clearest on this specific aspect.
On a broader level, I will reiterate my suggestions on Phabricator on working with the community:

[...] the scope and timeline of any planned community rollout, and the importance of working with the editor community and respecting a future consensus on whether to deploy it, should be made as clear as possible to avoid a repeat of Simple Summaries.

Chaotic Enby (in solidarity · talk · contribs) 12:31, 16 August 2026 (UTC)Reply
I'm with the rest of the group in saying that the NPOV suggestions are high risk and difficult to get right. I'm excited about trying to figure out if a good suggestion model can be developed for simplifying language. There is a large body of research showing that Wikipedia's science content is often way and way too complicated and linguistic complexity is an important part of that. We've had discussions on Wikipedia about using heuristics (sentence lenght, number of syllables per word), which is contentious. Perhaps an AI model can do this better than a heuristic, as it can distinguish between a long sentence with an easy sentence structure and an impenetrable long sentence. In solidarity, —Femke (talk) 🐦 15:27, 16 August 2026 (UTC)Reply
@Femke I have tried both AI models and conventional readability tests like Flesch–Kincaid and this is an area in which current AI tech cannot and should not be used. And my conclusion was that conventional readability tests also suck.
NPOV is another area where AIs cannot and should not be used.
We can use AI to find typos, but fixing them actually requires an experienced Wikipedian. Polygnotus (talk) 15:33, 16 August 2026 (UTC)Reply
I too have plenty of experience using AI for this purpose and find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. As these edit checks are only about identifying overly complicated text and because we have an abundance of low-hanging fruit in this area, I'm confident that a tool can be developed. The big question is if within the limits of a certain budget, it can become good enough, and to what extent it attracts the right editors to fix it. In solidarity, —Femke (talk) 🐦 15:40, 16 August 2026 (UTC)Reply
One of my common complaints at FAC (especially for scientific articles) is choppy writing style (sort of the opposite problem from overly difficult text). I often advise authors that they should try using some of the LLM tools to make suggestions for how their writing could be improved. I don't know how well this advice is received. I also don't know how much it is used in the intended way of offering suggestions vs just copy-pasting the output into their article; that would be an abuse of the tool, but people abuse tools all the time and the fact that they do so doesn't mean the tool is bad. RoySmith (talk) 15:43, 16 August 2026 (UTC)Reply
find that they can be used successfully in identifying overly difficult text and decently enough for suggesting alternatives if one pays attention to subtle changes in meaning. Subtle changes in meaning convert text supported by reliable source(s) to text not supported by reliable source(s).
I'm confident that a tool can be developed Yeah, developing a tool is the easy part. The hard part is ensuring it is a net positive.
The big question is if within the limits of a certain budget The WMF is not the right party to create such a tool because they are unfamiliar with the challenges Wikipedia editors face and because they have a tendency to waste a lot of time and effort on creating shiny new tools while neglecting decades of tech debt.
The good news is that volunteer devs can do it, but they will run into the problem that it won't work as I explained above. Polygnotus (talk) 15:46, 16 August 2026 (UTC)Reply
Some teams are unfamiliar with the editing challenges. Other teams have a track record of listening, or have editors in their team. The editing team is an example of the latter. For instance, when people at enwiki and dewiki asked for better VE performance recently, the team managed to fix this tech debt really rapidly.
I'm very willing to help test this part of the tool. I have experience with editors doing this well and less well. In solidarity, —Femke (talk) 🐦 15:58, 16 August 2026 (UTC)Reply
@Femke You can try WP:SCRIPTREQ, the people there are a lot more agile than the WMF is.
Please ping me if you have something I may be able to help test it. Polygnotus (talk) 16:02, 16 August 2026 (UTC)Reply
Tbh I still think we should look at incorporating SEWP articles here as 'simple summaries', but that's a different discussion Kowal2701 (talk, contribs) 15:47, 16 August 2026 (UTC)Reply
I wouldn't assume they'd want us to. Simplewiki is a tiny wiki with just a few active people. If we incorporate their content their vandalism rates will go through the roof. Polygnotus (talk) 15:49, 16 August 2026 (UTC)Reply
They'd also get more good-faith editors though! What I'm thinking is that SEWP still exists as a separate wiki, we just have an opt-in feature for displaying one of their articles behind a button. IIRC @Ferien was sort of open to the general idea (may be wrong) Kowal2701 (talk, contribs) 15:53, 16 August 2026 (UTC)Reply
Adding a CTA, if we get consent from the Simplewiki regulars, may be a good idea. Polygnotus (talk) 16:04, 16 August 2026 (UTC)Reply
Yeah, I am personally quite open to the idea, though I am not too sure what our community as a whole would make of it at the minute. --Ferien (talk) 21:04, 17 August 2026 (UTC)Reply
What kind of typos is AI fit to resolve that WP:AWB is not? Czarking0 (talk) 20:13, 16 August 2026 (UTC)Reply
One thing I'm having in mind are typos that depend on the context to make sense of them (or to whether there is a typo to begin with). Chaotic Enby (in solidarity · talk · contribs) 20:19, 16 August 2026 (UTC)Reply
FWIW, the fixes in Special:Diff/1369578064 were all suggested by Claude. I imagine most of them would have been caught by other tools, but I was impressed by the flagging of Freilberg. I don't remember exactly what it said, but the gist was that it spotted that I had "Peter Freiberg" in one place and "Peter Freilberg" in another. It figured out that these were probably referring to the same person and while it didn't know which was wrong, it assumed one of them was, and left it up to me to figure out which. RoySmith (talk) 20:24, 16 August 2026 (UTC)Reply
And I note the firsts of those fixes is changing a direct quote in a way that is inconsistent with the source. Sure, you can probably justify that (MOS:QUOTE allows minor typographic things to be silently corrected) but I don't think that's something an AI should be recommending in any way. * Pppery * (alt) in solidarity 20:49, 16 August 2026 (UTC)Reply
So, you're saying the fix is correct, and if a human had suggested it you would agree with it, but since an AI suggested it there's a problem? RoySmith (talk) 20:56, 16 August 2026 (UTC)Reply
I'm saying it might be correct (not that it is correct), but determining whether it is is a judgement call I don't want an AI to make. * Pppery * (alt) in solidarity 21:22, 16 August 2026 (UTC)Reply
The AI didn't make the judgment call. It just brought this to my attention and I made the judgement call. That's why my name is on the diff. I'm really not seeing the issue here. I made an error, a tool alerted me to it, and I fixed it. How is this a problem? RoySmith (talk) 21:36, 16 August 2026 (UTC)Reply
The funny thing about using LLMs here is that they start working against each other. The default behavior of AI trying to "rewrite articles into formal encyclopedic tone" is to undo any text simplification (and to do so poorly, for instance replace "is" with "serves as", "uses" with "utilizes", etc). The "text simplification" category here, however, seems to be basically a catch-all. "Simplification" suggestions here range from stuff like Correct the misspelling of the player's first name from "Russel" to the proper "Russell" (which isn't even correct!) to "break up these sentences."
There's also the problem that telling someone to Rewrite the sentence for clarity and smoother flow is not actionable to the majority of people: if someone doesn't know how to copyedit then they don't know how to do that, and if someone does know how to copyedit they don't need those vague directions. It's like prompting people as AIs. Gnomingstuff (talk) 18:23, 16 August 2026 (UTC)Reply
(Ironically, the LLM prompt that judged suggestions reads in part You are a strict reviewer. Your job is to find flaws, not to be nice. Really weird feeling to be envious of an LLM, its opinion certainly seems to be taken more seriously and its tone is given much more leeway.) Gnomingstuff (talk) 20:42, 16 August 2026 (UTC)Reply
Who on earth outside the group of people responsible for this is going to read Gnomingstuff's report and consider those suggestions to be "reasonably accurate"?  Hex talk 16:02, 16 August 2026 (UTC)Reply
Pre-alpha or not, the WMF shouldn't be experimenting with ways to funnel freeform model suggestions to editors to begin with. Machine models should never be allowed to so directly influence the contents of the project, as even if the suggestions are individually found to be valid and reliable, there will still exist overall biases. An LLM will favor certain sources, certain topics, certain sides. This will be reflected in what suggestions are and are not made, and editors evaluating and implementing individual suggestions will be entirely blind to any larger systemic issues they would be enabling.
Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact, and on a group level by the diversity of our views. A monolithic model has neither, it predicts tokens. fifteen thousand two hundred twenty four (talk) 21:45, 16 August 2026 (UTC)Reply
"Humans have issues with bias too of course, but this can be counteracted on an individual level by self-awareness of this fact - Citation needed: "making people aware of their bias doesn't do anything to mitigate it." Levivich (talk) 22:45, 16 August 2026 (UTC)Reply
Citation: I made it up from first principals and subjective experience, the preprint will be out soon.[Humor]
I feel entirely comfortable claiming that editors who operate with awareness of their own potential biases will take steps to mitigate them in this structured environment where WP:NPOV serves as a strong guiding force. If you find this unpersuasive, so be it, it is human to disagree. fifteen thousand two hundred twenty four (talk) 23:58, 16 August 2026 (UTC)Reply
The operative word here is can. As human beings that are capable of independent thought and self awareness, we can choose to examine our own biases and seek out information and experiences to change how we think about the world around us and the assumptions we make. Large language models are capable of exactly none of those things, because of the very simple fact that they are computer programs. Claude is exactly as capable of herself awareness as MS Paint is of having an independent thought.
Not every person makes the conscious choice of examining their baises, or is even fortunate enough to exist in a socioeconomic situation to even be able to, but that is beside the point gurkubondinn 00:26, 17 August 2026 (UTC)Reply
"Research from Harvard found that the effects of personal interventions such as awareness raising at a personal level are positive, but short-lived."
"And the worst method, the one that actually has no effect at all, is to tell people to be good people, to be egalitarian, and so on. ... It is easy in the sense that it last for a short period of time, but it won't last very long. ... Now, when young people encounter this result, when they see that, yes, they were able to make change, but the change doesn't last, they get very sad, because they want a better world. And I'm not at all sad about that. ... So our brains change, our minds change, associations move around, but they always will gravitate to whatever is your cultural default. And so to bring about actual change, society around us has to change, and then we will move, and then the default will be a new default."
LLMs are biased because humans are biased. Because they're trained by humans, and humans train their biases into the machines. We are no more able to cure bias in machines than we are able to cure it in ourselves. There are plenty of ways in which human intelligence is superior to machine intelligence, but lack of bias isn't one of them (in either direction). Levivich (talk) 02:57, 17 August 2026 (UTC)Reply
Absolutely, I agree with you. That's why the models can never be "neutral" or "unbiased". My point was just that I think you misunderstood 15224's reply, people have the capability to examine and be aware of their biases, but computer programs don't. It's not automatic and it doesn't happen for all people (for various reasons), but we have the ability to examine our own biases because (unlike computer programs) we are conscious beings. It's a bad comparison is what I'm saying, and it anthropomorphises computer programs (that were created by biased people). I have a beef with whoever it was that decided to wrap LLMs in chatbot interfaces. gurkubondinn 11:06, 17 August 2026 (UTC)Reply
We're looking forward to getting more deeply into this with you all this week. Before that, I wanted to express gratitude to y'all for the perspectives you are sharing. From the concerns about transparency and community control over models of this sort to issues with specific suggestions you're encountering, please keep the feedback/questions/concerns/etc. coming.
And for anyone who is interested in seeing what these suggestions look like in practice, please do the following...
TRYING EXPERIMENTAL SUGGESTIONS
  1. Ensure you have the Suggestion Mode beta feature enabled
  2. Enable "experimental" suggestions by pasting the following snippet into your common.js: mw.loader.load( 'https://meta.wikimedia.org/w/index.php?title=User:DLynch_(WMF)/alwaysbesuggesting.js&action=raw&ctype=text/javascript' );
  3. Open a page in VE that has one of the 30,000 suggestions available. E.g. https://en.wikipedia.org/w/index.php?title=List_of_examples_of_Stigler%27s_law&veaction=edit
Note: this week, I'm going to see if we can share a spreadsheet so you can see all of the model-generated suggestions in one place rather than having to tap around the wiki looking for them. PPelberg (WMF) (talk) 00:59, 17 August 2026 (UTC)Reply
I wrote a small script so you can view the list of suggested edits on an article without needing to enable the beta feature, switch to VisualEditor, or add that line to common.js to enable "experimental" suggestions: User:DVRTed/sandbox/edit-suggestions.js. — DVRTed (Talk) 03:06, 17 August 2026 (UTC)Reply
The spreadsheet would be helpful, if only because I don't know which of the csvs is the "real" one. In general I think this kind of thing is much easier to review in spreadsheet form than article-by-article. Gnomingstuff (talk) 04:18, 17 August 2026 (UTC)Reply
@Gnomingstuff: what you described makes total sense to me.[i][ii] I've checked in with engineering and it turns out that compiling this list will take a bit of time. Assuming nothing unexpected turns up, you can expect me to return here with a link to a CSV you (and everyone else here!) can review before this week is over.
---
i. Knowing, definitively, what suggestions we need y'alls expertise in reviewing and being able to differentiate those from earlier iterations that we've since discarded.
ii. Seeing all of the suggestions in one place so that you don't have to hunt them down yourself. PPelberg (WMF) (talk) 22:49, 17 August 2026 (UTC)Reply
Is the intent that the U/I would just present these suggestions to the user and let them edit the article themselves if they opt to accept the suggestion? Or is the idea to have a "Make this edit" button that the user could just click? The reason I ask is that if it's the later, it would make sense to include some machine-readable marker in the wikitext (I'm thinking an HTML comment) identifying the source of the inserted text. And/or have a log of such changes. The idea is to make it easier for somebody to audit the performance afterwards. RoySmith (talk) 23:04, 17 August 2026 (UTC)Reply
The latter would fail WP:NOLLM, so that would be a complete nonstarter. The former is not great either. gurkubondinn 23:15, 17 August 2026 (UTC)Reply
WP:NOLLM says "Editors are permitted to use LLMs to suggest corrections to their own writing, and to incorporate them after human review. This is limited to spelling, punctuation, capitalisation, grammar, and other simple mistakes." So this would indeed be a starter in those cases. RoySmith (talk) 23:22, 17 August 2026 (UTC)Reply
The suggestions that Gnomingstuff went through go far beyond the very narrow exception in NOLLM. And this exception only applies to [an efitor's] own writing. gurkubondinn 23:44, 17 August 2026 (UTC)Reply
The latter would be impossible in the current implementation, anyway, the LLM is not prompted to provide an actual change and the suggestions are usually just stuff like "The sentence is long, contains a grammatical error, and could be expressed more clearly." Gnomingstuff (talk) 05:09, 18 August 2026 (UTC)Reply
@RoySmith: Good question. If/when we (staff + volunteers) come to think these suggestions are reliable and useful, the intention would be to offer a suggestion that would contain the following:
1. A description of the issue and type of fix that is needed. Both of which will need to be generic enough to make sense across the contexts it might appear within while at the same time being concrete enough for the people encountering it to know how to start on the path of acting on it. So for the "Simplify language" suggestion, it might be something like, "Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language."
2. A link to the local policy/guideline the suggestion originates from. This serves both as an opportunity for people encountering a suggestion to learn more and also a way to ensure that suggestions are grounded in project consensuses and conventions.
3. Two actions: one action to Dismiss the suggestion and along with it, a way to express why someone has elected that choice. And a second action – maybe we'd label it Rewerite? – that when tapped would A) focus someone's cursor into the span of text the suggestion thinks there is an issue within and B) cause the article text in question to enter a "revising state" so people are clear about where exactly their focus is needed (see screenshot below). From there, the responsibility would be on the person acting on the suggestion to decide what to fix and how to fix it. Said another way: there are NO plans for these suggestions to make fixes with the click of a button let alone to describe specific solutions.
Screenshot showing the revising text state within Suggestion Mode
If you'd appreciate a more concise answer, what @Gnomingstuff described here is spot-on. PPelberg (WMF) (talk) 05:32, 18 August 2026 (UTC)Reply
@PPelberg (WMF): So how would you cram enough information in such a tiny area to give the person all the information they need? Are you aware that many guidelines are like 7k words? If you give a newcomer a link to WP:NPOV, that is obviously not enough to have them actually be able to judge the neutrality of an article if they do not have relevant experience and knowledge of the field. And what will happen when inevitably people complain that edits are not improvements? Will this just WP:BITE newcomers even more? Polygnotus (talk) 05:40, 18 August 2026 (UTC)Reply
@Polygnotus: great spot. The questions you're asking sit at the very core of this project.[i]
We seem to be aligned in thinking[ii] that some suggestions may be more harmful than helpful to show to newcomers. They're likely, as you alluded to, too complex and experience-dependent to distill down into a relatively small piece of guidance. If we don't account for this, we could lead newer folks into publishing edits that experienced volunteers revert or respond to with hostility. This could in turn drive these potential contributors away.
And while I don't think we can know for certain which suggestions those are in advance, I think we (staff and volunteers) are develpoiong a pretty good sense that conversations like this one are helping us refine.
I also think we'll learn a lot over time by trying things out, and there are a few features we've put in place to help:
Experimental suggestions: suggestions can be enabled as either default-on or experimental. The latter means that people who have explicitly enabled the soon-to-be-available setting have the space to safely experiment with suggestions. Through that testing, they can decide who—if anyone—an experimental suggestion should be shown to by default.
Tags: all edits in which someone sees and/or acts on a suggestion are tagged. You can see these in action by filtering Special:RecentChanges for Edit Suggestion seen or Edit Suggestion used.
On-wiki configuration: volunteers can independently decide the minimum number of edits someone must have published in order to see a suggestion. If you visit Special:EditChecks and look at the link suggestion, you'll see that volunteers have set the minimumEditCount value to 1000. This means the suggestion will only be shown to editors who have made ≥1,000 edits.
The idea is that, together, the above will let us see the kinds of edits these suggestions lead folks to make in practice and, with that, decide whether, how, where, and to whom they're shown.
How does this sound to you? What, if anything, do you think we might be missing or misunderstanding?
---
i. I hear the questions you're asking as something like: "How might we translate a great deal of nuance and complexity into a format that is simultaneously 1) succinct and simple enough that newcomers will engage with it and 2) explanatory enough that in doing so newcomers will be equipped with the information and know-how they need to act on them in ways experienced volunteers see as constructive?"
ii. Please correct me if I've misinterpreted what you've said PPelberg (WMF) (talk) 05:32, 19 August 2026 (UTC)Reply

I'm pessimistic about this, but I have a very high bar for pessimism/hopelessness to stop me from supporting a low-stakes experiment. Giving this tool to experienced users to try out seems like one of those low-stakes experiments worth trying, in case there's a way to make it work (again, I'm pessimistic, but possibly with certain constraints regarding task and topic?).
But also, just to put a finer point on something, because I think it does good to repeat it now and then: there's the worry about the quality of these suggestions, but there's also the worry about, for lack of a better word, branding. The branding that the WMF has begun using, "knowledge is human", is an echo of a popular sentiment here. At a time when absolutely every company and every project is cramming in as much AI as possible, Wikipedia is mostly headed in the other direction. That's a good thing IMO, but as a result, any pitch someone has for an LLM-based tool is evaluated with a handicap applied. If you'd otherwise be graded on a scale of 1-10 where 1 is harmful and 10 is helpful, start by subtracting 2 or 3 for "we don't want to be associated with that" (or, alternatively, "we see these as detrimental by default"), and it needs to be really useful to get a critical mass of people behind it. But no complex tool starts its life as an 8+ on that scale, I don't think, so super-low-stakes tests are IMO a good way to start working towards it. FWIW. Rhododendrites talk \\ 22:06, 16 August 2026 (UTC)Reply

I really appreciate and agree with the branding note Rhododendrite gives here. Best, Barkeep49 (talk) 22:48, 16 August 2026 (UTC)Reply

A humorous interlude

[edit]

I know people love to make fun of AI hallucinations, so I couldn't resist posting this fun map (4:38 in the video). Strange spellings aside, Hoboken, NJ has gotten transported to the midwest, Chicago is on the Pacific coast, and Boston has been relocated to Colorado. I want some of whatever it's smoking. RoySmith (talk) 02:12, 17 August 2026 (UTC)Reply

Also available in Africa. I think I'll stick to the Commons maps. Certes (talk) 14:30, 17 August 2026 (UTC)Reply
These LLM map-fails are legion. My take away from these examples is that they are not only vivid examples of LLM hallucination, and thus LLM unreliability, but they are also vivid examples of the unreliability of humans, because each one of these published hallucinations is only possible because humans obviously failed to check the work--to even look at the maps--before publishing. Humans are unreliable: an important lesson for any crowdsourced project. Levivich (talk) 14:57, 17 August 2026 (UTC)Reply
I don't know about this particular channel but loads of them are (almost) fully automated with no human in the loop. Polygnotus (talk) 15:39, 17 August 2026 (UTC)Reply
Then again, aside from quite a few mistakes, have you all seen the "Chloe vs History" channel on youtube? Quite the advancements in AI lately. Some of the advancements have been made on this channel, which should probably have a Wikipedia page. Randy Kryn (talk) 15:42, 17 August 2026 (UTC)Reply
That Africa one is from a US State Dept presentation. That ain't no YouTube clickbait, that's supposedly professional humans at work. You'd think they'd have looked at the slides before making their presentation. And, I'm speculating here, but I bet multiple humans were involved in that, because when the US State Dept gives a presentation at an int'l conference, I don't think it's just one person creating and making the presentation all by themselves. Maybe. But either way: evidence that at least sometimes, even professional humans presenting at an int'l conference obviously don't bother to check their work. I guess my point is: what makes LLMs unreliable isn't just that they hallucinate, it's that humans won't catch it because sometimes they don't even bother to look. Another well-known example is lawyers submitting briefs to courts with hallucinated citations -- literally licensed professionals in the performance of their profession. So this isn't a problem limited to youtube clickbaiters or "kids on the internet," even trained and licensed professionals, even on a global stage, succumb to laziness. At least the lawyers get fined for it -- now there's a new fundraising channel for the WMF: fine editors for publishing hallucinations on-wiki! Levivich (talk) 16:15, 17 August 2026 (UTC)Reply
At least with hallucinations, they're usually immediately obvious to anybody who bothers to look. This particular example is clearly YouTube clickbait. The bottom-feeders who produce these things don't give a whit about accuracy, just that they can churn out some mildly entertaining garbage videos that collect likes and shares and other revenue-producing metrics. But the fact that people are willing to abuse a tool for their own commercial benefit doesn't mean that the tool is inherently worthless. People abuse wikipedia in all sorts of ways (spam, SEO, reputation management, advertising, etc). Does that mean we should shut Wikipedia down? RoySmith (talk) 15:49, 17 August 2026 (UTC)Reply
PS, it doesn't take AI to generate garbage maps. Some people are able to do it all by themselves with nothing more technologically advanced than a sharpie. RoySmith (talk) 16:22, 17 August 2026 (UTC)Reply
Sure, and wikipedia had hoaxes and plenty of innocent misinformation long before LLMs became popular, but the problem, of course, is one of scale: LLMs make this sort of thing like 100x more common than before. It used to take hours to write a good hoax on wikipedia, now it takes minutes. Levivich (talk) 16:30, 17 August 2026 (UTC)Reply
"Chiiicago". AAND it's in multiple places at once. Ladies,Gentlemen and enbies, The Chicago hivemind. Starlet 01:43, 25 August 2026 (UTC)Reply

Back to business after the humorous interlude, a full ban on AI?

[edit]
  • This would be an overreaction. User:Polygnotus and many others have been building various LLM-powered tools, including ones that are used to detect LLM edits (User:Fermiboson/AIlog) and to fight vandalism (Wikishield - LLM is optional). There are more than one hundred editors using these tools. We do want WMF to experiment and implement the best ideas. Alaexis¿question? 07:06, 18 August 2026 (UTC)Reply
  • A full ban does not make sense. We already have a range of community tools that do cool things with Wikipedia using AI. In particular I want the best available tools for review, including those that take advantage of AI for trainable pattern matching and classification. That includes anything that helps with slop-detection, edit checks, citation checks, page linting, and vandal-fighting. This experiment falls under review... with a higher bar for quality I can see a range of suggestions being useful, particularly after a few rounds of focusing on quality.  SJ + 01:07, 19 August 2026 (UTC)Reply
  • I would support a ban on generative AI touching any Wiki content regardless of whether LLMs improve. JoelleJay (talk) 11:30, 22 August 2026 (UTC)Reply

Continued discussion

[edit]

Thank you all for keeping the conversation going, and thank you to Gnomingstuff and others for taking the time to look through outputs. @PPelberg (WMF), @SSalgaonkar-WMF, and I have read the whole thread and tried to pull out some of the most important things we heard and questions being asked. Peter and Sucheta, please add if I missed anything (well, for that matter -- anyone please tell me if I missed anything). I'm just going to list these questions/topics for now, and we'll keep working through them and hopefully you'll keep discussing with us (there are many of you and not as many of us!) This is to build on what I posted above and what Peter posted above.

Firstly, I think we're on the same page about something especially important: we're not going to deploy LLM-backed suggestions to anyone unless this community is supportive of it. And if they do become deployed, they'll be configurable via Special:EditChecks the way existing checks and suggestions are (i.e. community could decide who gets to see them, change what they say, what articles they show up on, or turn them off). I hope that in projects like these, we at WMF are bringing capabilities to volunteers so that they can produce the kinds of checks and suggestions that help both newcomers and experienced editors get the most good wiki work done with the least amount of drudgery. There is also a question of whether these suggestions would be a vehicle for content suggestions -- the answer there is no, we are not designing these to propose text to the user; rather they point out places the user should look and what they should look for. The LLM explanations that Gnomingstuff pointed out initially are in those files as a way for us to understand internally why the LLM is making the suggestion; they would not be shown to editors. (But I know there is a more subtle question here -- when just pointing out a sentence that needs attention in some way exerts that sort of influence).

In that vein, I wanted to say that this current project is more about figuring out what it might be like for checks/suggestions to be generated via LLMs, than about what those exact types of suggestions are. If NPOV is not a good one to pursue (volunteers here have given many good reasons why NPOV is particularly tricky), then we should put that one down in favor of simpler ones to try out. (Worth noting that checks and suggestions are being generated via a bunch of other ways, too, like simple rules, text match, simple models, etc.)

Also just a quick nomenclature thing:

  • Checks: this refers to edit checks that react to what the editor is doing at that moment in the editor. e.g. they paste in a blob of text from ChatGPT, and the check pops up right then and says, "Please avoid copying text from other sources".
  • Suggestions: this refers to suggestions for improvement that have been pre-calculated and are waiting to show up when someone clicks Edit, e.g. they open up the editor, and there is a box that says, "This link appears more than once in this section." Right now, Suggestion Mode is available on English Wikipedia as a beta feature, and suggestions are available in the feed on Special:Homepage.

You can see all the checks and suggestions and their statuses at Special:EditChecks (note that "experimental" checks are only available to people who have a specific user script installed).

Sorry -- this post is getting long. But on to the list of questions we want to be able to talk about here, raised in the thread above. This list of questions is not something that we just want to provide answers to and then expect that everyone will agree with us. This is a complex area, and we don't have all the answers, but we're trying to figure them out with you.

  1. What kind of communication with communities has there been on this project so far?
  2. Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing?
  3. How did we / do we QA lists of suggestions like these?
  4. What if communities don't want certain suggestions on their wiki?
  5. What determines whether suggestions are good enough to go beyond testing?
  6. Is NPOV appropriate to point an LLM at, given its nuance and complexity?
  7. If an LLM provides suggestions, how do we make sure that the editor isn't swayed/biased just by virtue of it coming from Wikipedia (a trusted source)?
  8. Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things?
  9. How could the models we use be transparent, open, and have community controls/auditing?

Alright -- more to come as we get into those questions. -- MMiller (WMF) (talk) 23:53, 17 August 2026 (UTC)Reply

@MMiller (WMF): Hi Marshall!
Because of a long history the relation between the community and the WMF is, let's say, far from perfect. And arguably worse than ever.
This is obviously not your fault, but its important to set the scene.
Because Wikipedians are smart and informed they are, generally, skeptical and wary of AI.
This is the correct position in 2026, because we do not yet know the impact of AI on the environment, jobs, and the upcoming resource wars.
The community generates the value and people think they donate to support it. The WMF is terrible at doing the community actually wants and needs. MediaWiki's tech debt is a sad joke. Community members who try to explain what we need from the WMF are routinely ignored. For decades.
Instead of working on the things that are important, the WMF nerds build shiny new toys. Debugging old code sucks. Doing something fun with AI/ML is fun. Because WMF leadership is terrible it seems (from the outside) that no one is working on the important stuff while we get an endless stream of halfbaked projects that waste a lot of time and money. The WMF basically never finishes a project it starts, overcommits at an early stage, and only gets feedback when its too late.
The WMF has a tendency to drop in and reveal they spent a lot of time working on a terrible idea, without asking input from the community, and then are surprised when the community rejects it.
The WMF is terrible at communicating its wishes and goals and what its working on.
The previous WMF attempt to "do something with AI" was a terrible idea and everyone hated it. It proved yet again that the WMF does not understand the Wikipedia community, what writing an encyclopedia means, and (the limitations of) AI.
The WMF did not learn from this and is again presenting a half-baked plan they spent a lot and time and resources on.
The community desperately wants and needs the WMF to succeed and produce great software at a reasonable cost, but that has yet to happen.
Quite a few members of the community are nerds who have spent a lot of time playing around with AI so they know what its limitations are in the context of writing an encyclopedia.
As the person behind the AI Proofreader, AI Source Verifier, AI Editsummary etc I am clearly not some anti-AI luddite.
The WMF is actively making the community more and more anti-AI, and to be honest rightly so. The WMF has no respect for the hard work of countless Wikipedians.
The fact that the WMF does things like have an AI check for NPOV problems shows that the WMF does not understand AI or its limitations or how to write an encyclopedia.
The community could greatly benefit from responsible AI use in a few specific tasks, whereby the human makes the final decision and is responsible for the edit. And the WMF makes it impossible for me to communicate that to people.
At this point, we both know its too late to listen to my feedback. This terrible idea will continue no matter what. The WMF has spent millions on AI related stuff and any benefits to the community were not proportional to the amount of money spent. It doesn't matter to the WMF; they had fun and can put something cool AI-related on their CV and move on.
The WMF has a toxic positivity problem wherein honesty and negative feedback is punished and ignored and all criticism must be hidden below 7 layers of a compliment sandwich. Even people far more diplomatic than I am just can't deal with all the corpo-speak and manipulation.
In an ideal world the WMF would stop what its doing and actually listen.
LLMs present an unique set of challenges and even opportunities and the WMF is fucking it up for everyone else and does not seem to understand how much damage they are causing and can cause in the future.
Is NPOV appropriate to point an LLM at, given its nuance and complexity? No, and that is a silly question. I don't teach my fish braille and ask it to rewrite the bible. Tools are useful in some contexts and bad in others.
What if communities don't want certain suggestions on their wiki? None of the proposed suggestions in its current form would be an improvement.
Does it make sense for the message to the world to be "Knowledge is Human", but we are also using AI for certain things? No, of course not.
How could the models we use be transparent, open, and have community controls/auditing? That is impossible unless you accept the models are terrible compared to competitors. There are no ethical LLMs available. And making something like that would be unethical and require unethical actions.
Why are we working on projects like this instead of more work on the backlog of bugs and small improvements for editing? Because the important work sucks and nerds think its fun to play around with AI. Also the WMF does not understand its role; it should serve and protect the community. In any well-run company you do maybe 90% important stuff and 10% fun stuff.
In the future, the community should be involved at the earliest opportunity, when brainstorming. The WMF should learn from its mistakes. Reflect on Simple Summaries. What went wrong, why, how to avoid that in the future? Allowing random newcomers to act on typofixes proposed by an AI will just make a lot of people very angry, sets newcomers up for failure and degrades the quality of Wikipedia. I use the opposite approach; my software helps experienced Wikipedians to fix typos, and an AI helps filter out things that aren't typos, and no one objects to that.
Polygnotus (talk) 00:41, 18 August 2026 (UTC)Reply
Please assume good faith. No one is doing this to have fun or put "something cool" on their CV and move on.
Wikimedia has actually spent a terrifyingly small amount on AI infra and tooling, which is part of the problem here: when an experiment is run, fast iteration on prompts, benchmarks, evals, and models isn't second nature. There are indeed ethical language models, as others note below, and better orchestration would help choose the best for a given task.  SJ + 16:02, 21 August 2026 (UTC)Reply
Marshall (and colleagues), I will say that from reading your post earlier, I think it reflects a good attitude for approaching the issue. (I'm not going to go down the rabbit hole of the difference between words and actions and trust.) I will also note that I have no idea how much say you have in what projects you work on and how far you take them. (Although if "Senior Director of Product" isn't just a bunch of fancy words, I would guess quite a bit.)
I clicked your link to see what's currently on Special:EditChecks, and I will say that almost all of them seem like really good ideas. None of them (the ones I like, at least), require any AI beyond some if statements in a trenchcoat.
In short: I think we could completely ignore all this AI stuff and you'd have some really great and helpful projects to work on. (Others have been mentioned above.)
I know working with such a large community can be hard. I would certainly like to think that we could give you some advice on how to help us. Ask . We only bite sometimes. Thanks for taking the time to listen. LittlePuppers (talk) 01:47, 18 August 2026 (UTC)Reply
That is an excellent and fitting username. Polygnotus (talk) 01:53, 18 August 2026 (UTC)Reply
One time I said that I didn't know what "Director of Product" means (I am not a native speaker) and the ex-CEO of the WMF started attacking and accusing me because she couldn't handle mild criticism. Polygnotus (talk) 03:11, 18 August 2026 (UTC)Reply
Thanks -- just to take this opportunity to shed a little light on what I do and how we're set up:
  • We're set up as a bunch of "cross-functional" product teams. Cross-functional means each team has a few engineers, an engineer manager, a designer, a product manager, a data analyst, a movement communications specialist (give or take -- sometimes a couple teams will share people in certain roles).
  • The product manager is responsible for setting the priorities of the team, what order to work on them, deciding what is in and out of scope, deciding whether the results of an experiment show that something is worth pursuing further. They do this in collaboration with their teammates, not unilaterally.
  • As a director of product, many of these product managers report up to me. When I started at the foundation, I was the PM of the Growth team, and I've been around for eight years and am now in a management role.
  • So in my role, I look across the various teams and play a role in setting the overall priorities -- like which various potential editing projects the various teams working on contributors should prioritize, what the readers teams should prioritize. I work with director counterparts in engineering and design to do this.
MMiller (WMF) (talk) 18:28, 18 August 2026 (UTC)Reply
Re 6, and to some extent 9: Does the WMF have a suitable operational definition of NPOV to measure against? As I understand it, transformer-based models are fairly common in sentiment analysis nowadays, but I don't know if pre-existing models (if any exist) would be well adapted to an encyclopedic context, and I don't think they would be sufficiently interpretable. In any case, I'd say that it is something that is strictly more difficult than tone check, as well as likely being considered higher risk by the community. I can't be certain of course, but I would expect for the community to grant social licence to a model for NPOV specifically would, at minimum, require much better interpretability than typical for current models. Whether the team believes they can achieve that is probably something best answered by the technical staff, the community can only inform as to acceptance criteria. Alpha3031 (tc) 05:00, 18 August 2026 (UTC)Reply
@Alpha3031 They just take off-the-shelf selfhosted open-weight models, which are way worse than Claude Fable 5 (and other mainstream commercial offerings) like gpt-oss:120b aya:35b / aya-expanse:32b, llama4:scout, qwen3:235b and qwen3:1.7b and then give it a presumably AI generated prompt containing:
 Role: You are an expert Wikipedia Copyeditor and Reviewer.
    Goal: Conduct a professional audit of an article's plaintext so it aligns with Wikipedia core content policies and Manual of Style.

    Context & Constraints:
    - Plaintext environment: non-prose elements may be stripped (infoboxes, references, templates, media, etc.).
    - Non-markup focus: do not suggest wiki-syntax, template, link-formatting, or reference-format edits.
    - Awareness of extraction gaps: if text appears truncated from extraction, do not flag it unless clearly an authorial issue.
    - Focus only on prose, terminology, neutrality, and information structure.
And
    - Neutrality: remove peacock terms, bias, and unnecessary loaded language.
You appear to be overestimating the sophistication of their approach by a rather wide margin.
Polygnotus (talk) 05:14, 18 August 2026 (UTC)Reply
I am aware of that, given that I had read (and in fact posted the link to) said prompts above. Though, I wouldn't say hosted open weights models are necessarily "way worse" than commercial offerings currently given recent and especially upcoming releases (GLM at 700B especially is likely a lot easier to host than the 2.xT models, though still harder than 120B of course). The page does indicate that the team recognises some situations may require bespoke models though. Alpha3031 (tc) 07:05, 18 August 2026 (UTC)Reply
Open-weights models can be perfectly adequate for some use cases. For source verification the very modest gpt-oss-20b model worked as well as Sonnet 5. Detecting NPOV issues is just a much harder problem, and it needs much more context and probably better models as well. Alaexis¿question? 07:18, 18 August 2026 (UTC)Reply
I don't disagree that open-weights models (even older or smaller ones) can be adequate for many tasks. I just also wanted to point out that as of the time of the current discussion, there are several open-weights models in the 700 billion to 3 trillion parameter range that are comparable to the frontier in a much wider selection of tasks, namely Kimi, Qwen and the upcoming GLM 5.3 (which being much smaller, is probably going to be cheaper to evaluate).
Incidentally, it does appear that previous WMF research findings (m:Research:Test External AI Models for Integration into the Wikimedia Ecosystem#Evaluation Results and Findings) have pointed out the difficulty of NPOV assessment:

Some policy detection tasks are hard for humans and machines alike. Specifically for NPOV, precision across the three model families is very low. Models overall do better when detecting Peacock behavior. This is because while understanding neutrality requires in-depth reasoning, [bold mine] peacock behavior can be detected via language features.

so it's somewhat odd that they've picked it as a easy, low-risk task in this specific project.
I would say that adequate NPOV assessment is probably the hardest possible task to set as a goal, given that it requires the aforementioned tone check, a comparison with the sources, and also some sort of test to ensure the sources the other models see are an accurate reflection of the body of published RS more generally. It's definitely not suitable for a project intended to explore what can be done with relatively generic models. I think there could also be room to explore, e.g., faster community feedback cycles so that the WMF doesn't feel like it needs a whole year to develop a single task specific model. Alpha3031 (tc) 08:56, 18 August 2026 (UTC)Reply
Agree that NPOV assessment is the hardest possible task, also for humans. More importantly, the assessment relies on consensus. There is no authority that has the truth about whether something is NPOV. The whole point is that the community reaches a consensus looking at all available reliable sources. Delegating this to an LLM (with the mark of approval of the WMF itself, thus giving it an undue resemblance of authority) defies the premise of Wikipedia, which is based on decisions taken by consensus. Aggravated by the fact that LLMs are not at all unbiased by any definition of the term (and are often just plainly wrong). Ita140188 (talk) 10:27, 18 August 2026 (UTC)Reply
Agreed, NPOV is hard! Alaexis¿question? 12:52, 18 August 2026 (UTC)Reply
Thanks for the response, this is already much more transparency than the last time around and I appreciate it.
I am a bit confused as to whether these questions are for us and for you. Gnomingstuff (talk) 17:35, 18 August 2026 (UTC)Reply
@Gnomingstuff: oh, good clarification. The questions Marshall posted are a first pass at translating what we're hearing from y'all into a set of questions that we (staff) can then respond to one-by-one. Of course, if you think there are questions we've missed and/or questions that we've misinterpreted, please comment as much. PPelberg (WMF) (talk) 20:05, 18 August 2026 (UTC)Reply
I thought I would ring in to point to another AI-backed project WMF is doing and share our experience working with volunteers on it, given the interest here in this sort of thing. For context, I lead the Product Safety and Integrity (PSI) team here at WMF, we build security and safety features.
One thing we're working on is detecting abusive content using LLM-based models. Specifically to enwiki, last week we deployed a new feature that is only visible to oversighters, at Special:AbuseReview, which uses an open-weight model called CoPE to flag edits that probably need to be suppressed due to containing personal information.
This on-wiki feature is now being used enwiki oversighters to review and take action on what it raises. They tell us the practical accuracy rate they experience in reviewing the output is better than what they see in the reports they get from human users. And, when we were testing the model output in June and July, in batches through an off-wiki process, most of the true positives it was catching were not being caught organically on-wiki.
One thing I want to mention is the iteration and volunteer collaboration that it took for us to make something deployable. There was some initial volunteer skepticism, and we needed to demonstrate not only that it was valuable, but that this was going to be accurate enough to not waste precious volunteer time. That took some internal experimentation on our part to get something that showed enough promise on both those fronts that we could ask for support doing a round of manual labeling -- which is what really pushed it over the edge of accuracy to be a deployable feature.
This collaboration has helped us do something meaningful about doxxing on English Wikipedia over the last few months, that feels good to both WMF and volunteers. That is possible in large part because volunteers were willing to give us space to operate, and to keep an open mind that could be convinced by data. We also needed to be flexible in our own thinking and incorporate volunteer ideas, something PSI has gotten used to doing in our work. EMill-WMF (talk) 21:17, 19 August 2026 (UTC)Reply
Endorsing Eric's statement, as one of the oversighters involved in the testing/iteration he describes. This is catching OS-level material we were not catching otherwise, and is a clear benefit to the encyclopedia. In solidarity, asilvering (talk) 21:54, 19 August 2026 (UTC)Reply
I think there is a role for AI to play on Wikipedia and it's going to be in something kind of like this. Something Eric alluded to but doesn't say directly which I think is important: the first draft the WMF showed us was nowhere close to ready. From that the learning was that we needed to go even farther in how confident . The second draft - which is when we started doing manual labeling - was still not anything which would have been appropriate for use. It did lead to a learning that one element - personal information - was more reliable than the other OS criteria which got us to the third draft and the efforts to then backfill June & July. That got us a lot of incredibly sensitive information that wasn't getting reported and which is now getting appropriately oversighted. I was the one who ran some data after we completed June to compare the true positive rate against email reports and found it to be in the range of 5% higher. I plan to revisit these stats soon and my hope would be that we'd have a higher difference between email reports and these flags for two reasons: 1) the classifier has been improved since I ran those numbers originally 2) some of the "obviously needs OS" tickets we'd have gotten in the past, we won't be getting because OS will be oversighting edits before some other qualified person finds them and reports them. I am really glad we're getting these signals now to stop some very sensitive personal information from being exposed against policy, we wouldn't have been able to do it without the AI classifer model, and it did take an iterative process to get there. Best, Barkeep49 (talk) 23:05, 19 August 2026 (UTC)Reply

Downloaded the most recent csv, picked an NPOV one at random:

  • "1358549724,The_Dying_Rooms,10791564,NPOV,Remove biased and profane language describing the Chinese government and replace it with a neutral summary of the government's response.,"In the film, Blewett and others travel to mainland China to visit orphanages housing children abandoned due to the ""one-child policy"". The filmmakers stated that unwanted female and disabled children were left to die of neglect, allowing parents to have another child. Showing that China government is a piece of sh!t for letting this happened and being a coward by lying to our faces that it didn't happened and said that all the footage is fabricated to destroy the reputation of the china government.>",enwiki,df59b690-6c8e-4827-b809-76e90d409f9f,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

The statement the LLM objects to was in the article for less than a minute and long reverted by the time the suggestion was created. So one can add to all the above problems and errors that it also wastes times and resources by not checking the current version but some snapshot.

Profanity checks in general are a bad idea.

  • "1358746317,2024_G20_Rio_de_Janeiro_summit,72163669,NPOV,Rephrase the incident with the first lady using neutral language and omit the explicit profanity.,"During a speech about fake news, Rosângela Lula da Silva, first lady of Brazil, swore at Elon Musk, saying: ""I'm not afraid of you. Fuck you, Elon Musk"" (Eu não tenho medo de você. Inclusive, fuck you, Elon Musk, in Portuguese).",enwiki,fba2904a-9276-47e3-8f5c-27fe5aff71f6,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

No, we are not going to "omit the explicit profanity" from a quote, and suggesting things like this is a very, very bad idea.

  • "1357774252,God_Emperor_Trump,74631506,NPOV,Remove profanity from the description of the phrase on the sword to maintain a neutral and encyclopedic tone.,"According to Fabrizio, the phrase could mean 'here's your fucking tariffs'.",enwiki,0f886b1f-2a4c-4303-b06e-63da6700b71e,"Editors often revise this kind of wording, saying the tone is unbalanced. You can help rewriting it using a [neutral point of view](https://en.wikipedia.org/wiki/Wikipedia:Neutral_point_of_view).",Revise tone"

Again, it's a quote, from the creator of the sculpture. The AI should not make statements like "Editors often revise this kind of wording, saying the tone is unbalanced. ", which only works to influence newbies by making a false claim to authority.

And then there the internally contradictory advices, indicating the inherent stupidity of LLMs.

  • "1356142445,Allan_Segura_(model),80130058,simplify_language,Combine the two sentences about sexual orientation and activism into one concise sentence.,Allan Segura is openly gay. He is a transgender rights activist.,enwiki,d28d67bc-2b32-4cbb-9135-ed391c91c13b,Readers might find this text difficult to understand. Try rewriting this using shorter sentences and plain language. [Learn more](https://en.wikipedia.org/wiki/Wikipedia:Manual_of_Style#Vocabulary).,Simplify language"

So do we need to combine these two (very short) sentences into one sentence, or do we need to rewrite this using shorter sentences? Or, just perhaps, our readers are perfectly capable of understanding these two sentences and won't "find this text difficult to understand". What a joke. Fram (talk) 09:01, 18 August 2026 (UTC)Reply

@Fram Also, wasn't the fact that the WMF didn't do content the thing protecting them in lawsuits?
Let's say an article contains negative information about a rich person. If the WMF starts to mess with content, do they not open themselves up to be forced to make changes? I am not a lawyer. Polygnotus (talk) 12:23, 18 August 2026 (UTC)Reply
All good reasons not to use AI for this purpose and probably not for any purpose. Once again, the WMF is wasting what remains of its valuable technical staff (and its ample funds) on trying to drag Wikipedia in completely the wrong direction.
If the bot is criticising vandalism which was only live for a minute, I suspect that it may be reading page history (why?) rather than taking a snapshot. Certes (talk) 12:32, 18 August 2026 (UTC)Reply
A snapshot sample of a large number of articles will also catch revisions that lasted only for a minute. There are lots such edits that are reverted quickly.
Assuming that edit suggestions are not displayed when the content has changed, this particular suggestion would never have been shown - no harm down do anyone. Generating checks on a snapshot rather than live is an architectural decision driven by performance, complexity and other considerations. Alaexis¿question? 12:51, 18 August 2026 (UTC)Reply
Presumably that means any such system will need to be rerun on each relevant page each time they are edited, in case the edit touched the prompted part? Not sure how the current tasks handle this actually. CMD (talk) 12:59, 18 August 2026 (UTC)Reply
Not necessarily, it can be run once a month, with some kind of caching enabled not to re-check the vast majority of content that stays the same. Then the complex and time-consuming part (LLM calls) is done asynchronously, and the easy part (deterministically checking that the content stayed the same) is done when the user opens visual editor. Alaexis¿question? 14:34, 18 August 2026 (UTC)Reply
That is indeed how it's currently working. A large batch of the suggestions were pre-generated, and was poured into a fairly simple API that VisualEditor's suggestion mode knows how to query and match up to a document.
Presumably if this was a successful experiment that we decide together to scale up, it'd turn into a more complicated system where we do something like fire off a job that precomputes the suggestion for each new revision of a page. Still avoiding needing to do the expensive generation every time someone opens the editor, but keeping them more up-to-date. DLynch (WMF) (talk) 23:03, 19 August 2026 (UTC)Reply
Again, it's a quote, from the creator of the sculpture.
That's another LLM tic -- at least when it comes to Wikipedia edits/suggestions, they really hate direct quotes and will usually tell you to paraphrase them and/or do so themselves. (Example from a 2026 edit summary: "Paraphrased a lengthy direct quote regarding the skater's performance at the 2025 World Team Trophy into a concise summary. This improves readability and helps maintain a neutral, encyclopedic tone by removing overly detailed personal reflections.") Gnomingstuff (talk) 17:41, 18 August 2026 (UTC)Reply
The examples Fram shows above makes me more resolute in opposing this on principle. The WMF is trying to exert editorial control and hiding it by labelling them as "suggestions". JCW555 (talk)16:48, 18 August 2026 (UTC)Reply
@JCW555, I guarantee you they are not trying to exert editorial control with this feature. They're trying it because they think it will be helpful for beginner editors. We can tell them they're wrong and we don't like the feature without impugning their motives. In solidarity, asilvering (talk) 21:21, 19 August 2026 (UTC)Reply
But suggesting sentences should be reworded is the WMF trying to exert editorial control. To quote MMiller above "It would just point out the spot in the article that needs attention, e.g. "Does this sentence need to be rewritten to be easier to read?". Whether a sentence/passage needs to be reworded is a matter for the talk page amongst editors, not through the WMF via their AI. No reply from the WMF in this section has done anything to assuage that concern for me. If the WMF comes out and says that these "suggestions" wouldn't touch content at all, no matter how small or big, then I'd be a tad less aggressive in my opposition. JCW555 (talk)21:44, 19 August 2026 (UTC)Reply

Is NPOV appropriate to point an LLM at, given its nuance and complexity?

[edit]

Hi all, I'm Sucheta, product manager on the Machine Learning side of this work.

Reading everything you all have said, it’s clear we shouldn’t move any further with the LLM-generated NPOV suggestions. We're dropping that type rather than trying to iterate on it.

@Ita140188 said this really well: There is no authority that has the truth about whether something is NPOV. Makes sense to me – NPOV isn’t just about word choice; it's about the representation of a topic that editors decide on by weighing the available reliable sources against each other and reaching consensus. We had wanted to take a crack at it to see what the LLM would produce, so thank you for looking at these and thinking about them.

I do want to make the distinction between these LLM-backed NPOV suggestions and the Revise Tone suggestions that are in production now on this wiki. Those Revise Tone suggestions come from a model called BERT that we’ve fine-tuned for a narrower scope. The model is trained to notice peacock language, based on 20,000 examples of revisions where the "peacock" template was added or removed. We think these have worked out well, and thousands of them have been actioned by both new and experienced editors on this wiki. Having them available also made it more likely that a newcomer would make a constructive edit, than when they open the editor without some suggestion inside. You can see them on Special:Homepage in the suggested edits feed (if you check these out and have thoughts, please let us know).

What about the other types of LLM suggestions besides NPOV? Well, let’s keep talking here about whether they have potential. Note: we're working on making a spreadsheet available to you all so that you can see all of the suggestions in one place. -- SSalgaonkar-WMF (talk) 17:13, 18 August 2026 (UTC)Reply

Thanks for the response. Good to hear about the NPOV feature.
The main issue is mostly the same across the board: the actual LLM suggestions have issues as above, but including broad categories without the actual suggestions is just confusing and provides almost no context. This is going to affect any possible category, it's just structurally inherent to the task as I understand it. Other than that:
MOS:GEO: Two issues I can think of:
  • Seems near-certain to inadvertently wade into a geopolitical quagmire of some sort.
  • Less dramatically, a large proportion of these suggestions are really just suggestions to treat everything as first reference (the "Atlantic Ocean"/"Atlantic" thing mentioned above). I don't think there's any way to get around this while working at the single-sentence level.
Simplify language: Two issues again --
  • The text parsing needs to be fixed before anything is done with this since otherwise the suggestions won't make any sense, especially the issue of parsing multiple sentences as one (e.g. King's fourth novel, Euphoria (2014), was inspired by events in the life of anthropologist Margaret Mead. It won the inaugural Kirkus Prize for Fiction and the 2014 New England Book Award for Fiction, and was a finalist for the 2014 National Book Critics Circle Award. Euphoria was listed among The New York Times Book Review's 10 Best Books of 2014, TIME's Top 10 Fiction Books of 2014, and the Amazon Best Books of 2014. -- they aren't visible on-page but there are unicode separators between many of the clauses, maybe that's related).
  • The category scope seems to be off. The vast majority of suggestions are to break up sentences, comparably little about actually simplifying language -- if anything, the language suggestions I've found seem to be suggesting the opposite, to make language more complex. But suggestions for typo fixes etc. also show up here.
Gnomingstuff (talk) 17:58, 18 August 2026 (UTC)Reply
That unicode-characters thing is actually deliberate. The data comes pre-massaged into the exact form that works in VisualEditor's search (non-text gets replaced with a opening and closing internal model tag, so that  is actually <ref></ref>). E.g. If you go to the Lily_King page that quote's from and paste it into the VE search box, it should highlight that entire paragraph, covering the citations. As far as I know, that replacement got done as a post-processing phase after the initial suggestion-generation, though I wasn't directly involved so there's a chance I'm wrong.
...that said, you did make me realize that we're accidentally not comparing them in that form in our final "has the user already changed this bit of text since they started editing" check, so gerrit:1326927 will make all these suggestions that cover citations / templates actually visible for review. DLynch (WMF) (talk) 20:54, 18 August 2026 (UTC)Reply
  • Questions about peacock language; Does the function ignore direct quotes? Any suggestion to change the wording of a direct quote shouldn't happen. Also, can we can we get it to come down hard on subjective words like best while being less aggressive on words that might be verifiable facts like largest? And maybe even less aggressive for largest known and largest on record? --Guy Macon (talk) 18:39, 18 August 2026 (UTC)Reply
    @Guy Macon: Great question. This suggestion can be configured to ignore quoted content which, at present, is exactly what en.wiki has done.
    You'll notice that ignoreQuotedContent within Special:EditChecks#tone is set to true. If you'd like to learn more about exactly about how quoted content is detected, T414715 contains more details. Could you please let me know if anything you see (or don't see) brings other questions to mind?
    Now, to the second question you're asking...
    Assuming it's accurate for me to understand it as something like "How might we specify the suggestion based on specific words/phrases?" I wonder if you think TextMatch could be helpful here. In essence, it enables volunteers to write custom suggestions to appear when predefined words/phrases are detected with an article. PPelberg (WMF) (talk) 19:45, 18 August 2026 (UTC)Reply
An update on the NPOV suggestions. As of ~30 minutes ago, we've removed all NPOV suggestions from the experimental batch of model-generated suggestions. Thank you all for the quick feedback here and for trusting us to hear you.
Note: if, by chance, you happen to still encounter one, can you please let us know? For now, we implemented the above in a bit of a fragile way so we can get something out quickly and will come back to make this more robust in the coming days. PPelberg (WMF) (talk) 22:11, 18 August 2026 (UTC)Reply
Thanks, love this fast and encouraging response. If peacock language detection is working well with a training set of 20,000 examples, is compiling training data a helpful step for catching other narrow style issues?  SJ + 01:20, 19 August 2026 (UTC)Reply
Great question! I think so - though this project is meant to test how viable it is to generate suggestions without manually compiling the training data you described.
If we were to place the LLM-generated suggestions and Tone Check approaches on a spectrum, Tone Check would sit at one end: it works well for identifying a well-defined issue, but it took us over a year to build and deliver, with training and evaluation being two of the most time-consuming pieces. We're now considering three ways to generate suggestions using models, among the many other approaches we're exploring:
  • generic models with task-specific prompts (what we're testing here)
  • generic models plus fine-tuning
  • specialized, bespoke models
So the question we're really asking is what amount of rigor is required to produce suggestions that are sufficiently reliable and useful?
We're asking a similar question about evaluation: whether faster methods like LLM-as-a-judge can tell us if suggestions are reliable without hand-labeling a large test set. Even here we review a number of samples manually to make sure the judge is scoring appropriately; that sample is just much smaller. We also created an eval dataset of past user edits per suggestion type so we could compare the model-generated suggestions to actual edits.
What do you think about this lens? Do you think there are some suggestion types that could be supported by this approach of using generic models without fine-tuning? SSalgaonkar-WMF (talk) 15:42, 20 August 2026 (UTC)Reply
How much effort is something like this? ScottishFinnishRadish (talk) 16:01, 20 August 2026 (UTC)Reply
Thank you for sending this! I really like this idea of experimenting with LLMs to generate basic copyediting suggestions. Reiterating what you said: they seem to be relatively low-risk, simple, and as a result, something newcomers could handle. In terms of effort, I don't think it would require an impractical amount to build a dataset like this (as your demo even shows).
We started to explore copyediting suggestions using MoS guides around capitalization and grammar, but we didn’t feel confident enough about their quality and utility to include them in this initial dataset. The capitalization suggestions scored lower in our LLM-as-a-judge evaluation than other suggestion types, and our manual review showed that “grammar” was too broad a category; we couldn’t come up with one label or description to describe the range of issues surfaced by grammar suggestions.
These both feel like solvable problems, and I’d really love for us to take another look. Would you be willing to share the prompt you sent to ChatGPT? SSalgaonkar-WMF (talk) 13:54, 21 August 2026 (UTC)Reply
Find some typos that can be fixed or copy that can be edited at https://en.wikipedia.org/wiki/Mircea_II_of_Wallachia. If I were going to refine it I would probably break results down to things that need review from someone with varying levels of experience so you can filter the output to users based on experience. For instance, checking if the source said first born or firstborn is a great task for someone trying to step up from beginner editing. ScottishFinnishRadish (talk) 15:10, 21 August 2026 (UTC)Reply
I don't know if the technology can support this, but it would be nice if we could have multiple queues of suggestions in different areas. A queue for spelling and grammar fixes and a queue for source-to-text validation in STEM articles might appeal to different people looking for work. RoySmith (talk) 16:02, 21 August 2026 (UTC)Reply
When I was looking through the csv before I didn't find errors when the llm was identifying the wrong units being used or similar. In my personal testing I've found it good at picking up tense mismatches (occasional errors sure but overall picks things up I've missed when rewriting something). This sort of pattern fixing seems much less likely to cause an issue than setting an llm to gambol over fields of longer text, as well as being a simpler spot and fix for newcomers. CMD (talk) 00:45, 19 August 2026 (UTC)Reply
My apologies if I just read past it and didn't notice, but where is this CSV file? RoySmith (talk) 00:49, 19 August 2026 (UTC)Reply
@RoySmith In , GnomingStuff dug it up. CMD (talk) 01:11, 19 August 2026 (UTC)Reply
Got it, thanks. RoySmith (talk) 01:14, 19 August 2026 (UTC)Reply
We'll be sharing a clearer spreadsheet version (with the most up-to-date entries) tomorrow, per Peter above. HTH. Quiddity (WMF) (talk) 01:17, 19 August 2026 (UTC)Reply

Why are we working on suggestions?

[edit]

The first reason we're working on this is because of the need to get more new people involved in editing. Being a newcomer has always been hard, and is especially hard now that people spend most of their online time on mobile. When newcomers (especially on mobile) open up the editor for the first time, it is overwhelming and they often just leave -- they are like "Wow, scary, nevermind." (here are some interesting survey results about this moment) But we've seen that when we point out specific bits of the article that could use improvement, the newcomers are much more likely to do something constructive and stick around. We've tested many of the checks and suggestions in Special:EditChecks and they have had these measurable positive impacts. We’ve also been inspired by the tools/scripts/gadgets that volunteers have built that do similar things (some examples here).

So this project here (the LLM suggestions) is another way of learning how we might find more kinds of suggestions in the vein of "how can we help newcomers on mobile be more and more constructive and more likely to stick around?" (in ways that align with policies, values, and existing editor workflows).  It’s also worth noting that the newcomers who start with suggestions often wander off on their own in the wiki once they get comfortable. The design helps with that, because the suggestions happen inside the Visual Editor, i.e. you have to edit in VE to get them done (as opposed to, say, a separate interface).

Secondly, we also think that this can help lower patroller burdens at a time when those burdens are increasing because of AI slop, as Gnomingstuff has pointed out. (i.e. newcomers making constructive edits in the first place lowers burden on patrollers from newcomers being confused). For example, we are developing a way to deter and label edits when people are pasting content from an LLM.

And thirdly, we think that these suggestions can be helpful for experienced editors, too.  A few of you have said in this conversation that experienced editors don’t need help finding improvements to make, but we have also heard from many who appreciate suggestions like these. In my own editing experience, I often click edit to do a specific change, but then discover a few other small things to improve via the suggestions.  So we think there is also opportunity here to help experienced editors get more wiki work done with less seeking/searching/effort.  It becomes a question of which of these signals to present to which users in which places.

How does this all sound? It would be great to hear from anyone who has seen these suggestions in action with newcomers or has been using them.

-- MMiller (WMF) (talk) 19:49, 18 August 2026 (UTC)Reply

Around here, the city is running an e-scooter pilot. Install the app, hop on a scooter, ride to where you're going. One of the interesting things they do is the app automatically imposes a lower speed limit on all new riders. Once you've ridden more than (IIRC) 10 hours, you get to go full speed. I think it also won't let newbies take out a scooter after sunset. I forget the details, but you get the idea.
I could see doing something similar here. Have some way of scoring suggestions for how risky they are. Fixing an obvious typo is pretty low risk. Rephrasing a statement that appears to be biased is higher risk. The type of article might also factor into it: an article about a WP:CTOP would probably not be the best choice for a newbie to learn on. The total neophyte would only get the safest suggestions. People who had gotten a bit more experience might be offered a wider range of suggestions. RoySmith (talk) 20:14, 18 August 2026 (UTC)Reply
MMiller (WMF), two thoughts come to mind:
  1. A good place to ask might be WT:AFC and similar (e.g. NPP)—on one hand, writing articles is kind of exactly what we don't want new editors to try, because it's really hard, but we see about every kind of possible error there: from formatting (a first heading duplicating the page title, formatting inside headings, malformed templates, all sorts of weird stuff) to tone (as established this is hard, but there are probably a few ways to detect COI) to referencing (citing Wikipedia, citing social media, just not referencing, weird formatting, duplicated refs). I see that some of it you have projects related to, but that's a handful more off the top of my head, and I'm sure people at those pages can think of more. I'd also be curious if you've investigated what the impacts of limiting this to visual editor are (i.e. how many new editors use VE).
  2. Another thing to consider is looking at what gadgets and user scripts are commonly installed. Your check about disambiguation links makes a lot of sense to me, because there's been a gadget which displays them in a different color for years. And some of those do make more sense as user scripts or being community maintained, but there are definitely some which would make more sense as part of MediaWiki or which could use some love. (Looking through my user scripts, there are a handful I don't really use, and some which are enwiki-specific, but also many which make a lot of sense to integrate or which I'm importing cross-wiki and haven't been updated for 8 years or something.)
I don't know if you're short on ideas or not, but those are what came to mind for me.
And I just reread your post and realized you already did #2. LittlePuppers (talk) 21:21, 18 August 2026 (UTC)Reply
@LittlePuppers: thank you for sharing ideas of places to look for and vet ideas for new Checks and Suggestions. I've added WT:AFC and Wikipedia:New pages patrol to to the MediaWiki page where we are bringing together this sort of information. If/when other ideas come to mind, we'd be thrilled if you'd add them directly. Of course, I'm happy to add them as well. Just give me a ping if/when something strikes you.
On the topic of gadgets and user scripts, we're with you here. In fact, doing what you described helped prompt the work we're partnering with @Alaexis on to turn a tool he wrote for identifying cases where a source might not support its associated claim into a new experimental suggestion. PPelberg (WMF) (talk) 23:13, 18 August 2026 (UTC)Reply
Re writing articles is kind of exactly what we don't want new editors to try, I'm not sure I agree with that. For some new users who don't know how to get started, having them fix typos might indeed be a good way to ease them into editing. But some people will already know what they want to write about. If you tell them, "No, that's too hard, we want you to fix typos for a while", all you're likely to have done is lost an opportunity to get a new editor hooked on the project. The very first edit I ever made was to create City Island Bridge. RoySmith (talk) 01:32, 19 August 2026 (UTC)Reply
Yeah, I was wondering if someone would push back on that. My wording there was probably too strong, main point being that it's really hard and, I suspect, very often discouraging. But then again, I do see, on occasion, someone who will read through guidelines and knows how to write and all that and write pretty decent articles very early on. To be honest, those are probably also the people who write decent articles later on.
Today, your first edit would have to go through AfC, and get declined for being unsourced... ah, those were simpler times. There was probably also more low-hanging fruit then. I have no idea where I'm going with this comment. I'm all for developing features to help out those who start in all sorts of ways. LittlePuppers (talk) 01:59, 19 August 2026 (UTC)Reply
The difficulty is that so very many new editors know what they want to write about, but what they want to write about isn't going to make it to article form at the present time. They want to write about themselves, their companies, their favourite YouTuber, the assignment they've been given, the really fun thing they made up...
Yes, we would lose them if we told them to fix typos. But we also lose them if we don't let them create that article they want, and we can't let them create that article they want. Many of them are also happily LLM-ing it up in their draft. I almost wonder whether there's a call for a service (human, AI/LLM, both?) that we try to push would-be article creators towards that assesses or helps them assess whether they actually have sources - and tells them to stop if they don't, or at least to try another subject instead. Perhaps an AI/LLM project could be given some basic guidelines about common problems (interviews are likely to be inappropriate; this source doesn't seem to be about the subject) and at least head off the ones that don't have a chance. A lot of people seem remarkably inclined to listen to what a machine says, possibly because it's seen as authoritative and knowledgeable on basically any subject. Meadowlark (talk) 06:27, 19 August 2026 (UTC)Reply
Hi @Meadowlark, I'm Rita Ho, director of design at WMF working as the design counterpart to Marshall with product teams working on these new editing tools. Your comment about helping editors to create new articles by providing more guidelines (like including sources) is related to another feature in development, called Article Guidance! This feature is aimed at helping newer editors who want to create articles to succeed. It's kind of like if Article Wizard could be tailored and offer specific support depending on the type of article being created (animal/building/person/etc). It includes initial source validation, notability risk assessment, and initial minimal content structure or "outline" for someone to get started, and is community configurable.
Sharing in case you and others are interested to learn about this other project and participating in testing and giving feedback. RHo (WMF) (talk) 22:01, 19 August 2026 (UTC)Reply
I do like having community-create outlines. That seems (at least in theory) to catch a lot of the types of mistakes I see with new editors (sources!). LittlePuppers (talk) 22:11, 19 August 2026 (UTC)Reply
It becomes a question of which of these signals to present to which users in which places.
Building on what Marshall shared above, I think it might be useful to consider that we're trying out a range of ways for generating the signals to power new edit checks and suggestions...
Some suggestions, like adding references and detecting when someone has pasted content from an LLM, are generated using deterministic rules.[i] Some look for the presence of specific templates, like citation needed. Others use small language models specially trained on Wikipedia edits to identify a specific type of issue, like finding issues with tone or suggesting images to add to articles. There are also suggestions written by volunteers using the TextMatch feature (inspired by AbuseFilter). We're also experimenting with a way for volunteers to create suggestions based on the presence of maintenance templates or missing template parameters.
I share all of this in an effort to communicate that we're eager and open to experiment with a signals from a variety of sources. What's most important to us is identifying ones that we collectively see as reliable and useful.
---
i. E.g. contents of clipboard metadata and the absence of a reference within a defined amount of new text PPelberg (WMF) (talk) 23:03, 18 August 2026 (UTC)Reply
Thanks, I've edited the mediawiki page to add two ideas, one to identify/highlight problematic text that a maintenance tag is referring to, one to point people to en:Help:Find sources if they try to use an unreliable source like a blog or social media. I think it's probably best if we focus on problems that would result in a revert for the sake of retention of newcomers (rather than minor stuff like MOS:GEO, where people can learn from someone editing their work). I'd also suggest working more closely with WP:AIT if you aren't already (if they have the time, people such as @Polygnotus, Alaexis (as I see you're doing), @Dreamyshade) as they'll likely have a lot of ideas and be able to discern some issues which may not be apparent. Otherwise I'd just encourage people to boldly edit the mediawiki page and add any ideas or whatever. Maybe it could have a second column for concerns about an idea? Kowal2701 (talk, contribs) 08:41, 19 August 2026 (UTC)Reply
Approaching this from the narrow goal of improving Wikipedia, I'm concerned to see newcomers and suggestions again mentioned in the same breath. Automatically produced suggestions need to be assessed individually by someone familiar with writing for Wikipedia, and that's not a newcomer. If the goal is not to improve Wikipedia but to increase the active editor count by making newcomers feel useful then this may be a viable scheme, in the same way that we reluctantly allow education projects to introduce so many errors to our articles. However, if that is the case then (yet again) editors and the WMF are pulling in different directions and we have a clear conflict of interests to resolve. Certes (talk) 09:18, 19 August 2026 (UTC)Reply
@Certes -- yeah, I understand. We've been talking about this since the early days of the Growth team in 2018: were suggested edits more about improving Wikipedia, or about retaining newcomers so that they could grow into editors who improve Wikipedia later? We generally erred on the side of retaining newcomers -- for instance, the first suggested edit was "add a link". Do the wikis really need lots more blue links between articles? Some wikis do, some not really. But it was a good task for newcomers to get their feet wet, have a succesful first experience, and want to come back again. We saw some newcomers "go on a run" where they did hundreds of those tasks over the course of several days. And we (happily) also saw many do a few of them and then go do some other, higher value edits on their own.
But the ideal is that we can do both at the same time: design tasks that are both healthy for newcomers and constructive for the wikis. I would say that "revise tone" is an example of this. I would be curious what you think of the diffs in this Recent Changes filter, which shows both links and tone diffs, highlighted by how experienced the person is.
And more generally, where you would come down on that trade-off between "invest in newcomers learning" versus "constructive edits now" (hoping, of course, that we could have our cake and eat it too with the right designs). MMiller (WMF) (talk) 21:28, 19 August 2026 (UTC)Reply
I rarely improve tone and am no expert on it but I looked at the first three recent changes by different editors. The first looks like a useful improvement from a new editor who clearly already has the skills we need. The second slightly misses the point: I've seen the film and the character's defining aspect is that he is a local legend. The third is also wide of the mark: much of the promotion is in the paragraph before the one the editor changed. Its edit summary of "changed tone(bot told me to)" is also concerning: perhaps the editor values obeying "the bot" above using their judgement or feels that this is what is required to become an accepted editor.
Having our cake and eating it would be nice, but I do tend towards constructive edits over using articles as a sandbox. Certes (talk) 21:46, 19 August 2026 (UTC)Reply
Taking a look at the tone edits, starting at the bottom:
  • Special:Diff/1370168493: The edit in isolation is ok but the edit summary suggests that it is AI-generated, and the many other edits they have pumped out such as Special:Diff/1369053156 and especially this promotional draft corroborate that. The ideal response here would be for them to stop using AI for promotional edits. I would also suspect an SPA based on this edit history.
  • Special:Diff/1370225916: Probably OK, no immediate red flags
  • Special:Diff/1370168654: Obvious promotional AI slop.
  • Special:Diff/1370168718: Probably OK, no immediate red flags
  • Special:Diff/1370168757: Obvious promotional AI slop by the same person who did the last one, which should illustrate the volume at which this adds AI slop to the wiki. (Also, the topic is potentially controversial.)
  • Special:Diff/1370169145: Obvious promotional AI slop by the same person, I'm going to skip their edits from here on out but just know that I am skipping a lot of bad edits as a result.
  • Special:Diff/1370170671: OK but not really a tone edit, and their edit history is somewhat suspect. Also, Special:Diff/1369948373 does not actually fix the tone, it just puts a band-aid over it, which is the other problem with these edits.
  • Special:Diff/1370171161: Probably OK.
  • Special:Diff/1370171537: Promotional AI slop (by someone else this time, Special:Diff/1350356412 is the smoking gun and they have several warnings on their talk page). As you can see from their edit history they are also pumping these out at high volume so I will also be skipping their many bad edits.
  • Special:Diff/1370171930: Grammar edit that does not actually fix the tone, it is still promotional
  • Special:Diff/1370172403: Not a tone issue. The fact that there is an obvious tone issue literally one word away does not speak highly of their competence in editing. (that is, competence in editing English-language writing, not Wikipedia specifically)
  • Special:Diff/1370174754: Probably OK, no red flags
  • Special:Diff/1370176147: Probably OK, no red flags
  • Special:Diff/1370176495: Obvious promotional AI slop and they didn't even bother removing the citation markers.
This is the thing that I have been trying to point out for several months now. The feature is a net negative. It may result in numbers going up in terms of edits by newcomers, but the workload for editors also goes up -- if they even notice it -- to a point where there are simply not enough people available to cleanup. This also means that revert rates are artificially low, because again, there are not enough people to even see them. Gnomingstuff (talk) 22:07, 19 August 2026 (UTC)Reply
Okay, well as a third person who has randomly gone through some of these:
  1. diff: change is an improvement; the paragraph is unsourced, and might be better removed. (My changes: pt 1, pt 2, pt 3; still room for improvement.)
  2. diff: may or may not be an improvement; I haven't seen the film, and I doubt the editor had either.
  3. diff: possibly worse, though I haven't seen the show.
  4. diff: likely improvement. Removes unsourced content.
  5. diff: minor but definite improvement. Edit summary includes "bot told me to". I'd probably also remove "even" or reword, but "one of the largest" is likely factual, albeit not sourced inline.
  6. diff: could be worded better but it's an improvement.
  7. diff: minor improvement, room for more work.
  8. diff: meh, at least there's a period at the end of the sentence now.
  9. diff: not really an improvement.
  10. diff: improvement.
A common theme is that many of these seem to involve people changing tone without the background knowledge to understand the article. "Tone" is also vastly oversimplifying the variety of problems these articles have. LittlePuppers (talk) 22:15, 19 August 2026 (UTC)Reply
I actually have seen the show; the former text was an accurate description of the character, especially given that the characters in Sunny are... extreme and exaggerated people. So this person and/or any hypothetical AI they are using does not know the difference between fictional characters and real people.
The other issue -- and this is an issue across the board with all newcomer tasks -- is that the template on the article is actually more specific than revising tone: it says that the article is written in a primarily in-universe style and should not be. I assume they were not told that in the task. Gnomingstuff (talk) 22:22, 19 August 2026 (UTC)Reply
Re: newcomers vs experienced editors this is partially covered by Peter's comment above. I.e. These features are entirely configurable (and extensible) by each local community, and importantly, that means that some of the types of Suggestion can be completely limited so they're only seen by highly-experienced editors. If you/anyone can think of types of Suggestions that would be widely useful for just experienced users, and if those Suggestions can be programmatically recognized by simple textmatching, or more complicated types of code within the extension code, or via a locally run/controlled LLM, then it should be possible to setup those kinds of things (to test, and if proven useful then to make available to all editors who fit the locally defined configuration).
Also, one of the main goals of the feature is to help provide editors (both newcomer and experienced editors) with a handy link to the relevant guideline/policy within the Suggestion card's text. That makes it easier for editors to learn (or remind ourselves) of the specific nuances involved in any particular fix. E.g. If someone is editing a disambig page, and they ignored the EditNotice reminding them of the basic guidelines, then when they add two links within a single entry (or stumble upon an existing entry which does so during their editing), there could be a Suggestion that informs them of that guideline and has a singular pointer to the details (versus the nine links in the current Enwiki Editnotice). I.e. Targeted micro-tutorials that show up when relevant. The main limitation for what can be written in the Suggestion card, is length, as it also needs to work in a mobile-sized screen. Quiddity (WMF) (talk) 22:53, 19 August 2026 (UTC)Reply
Just want to reiterate that I appreciate the response here, it is already a much better dialogue than before and that's what I was trying for, to keep the focus on the content and implementation details itself. It may surprise some people here but I am not blanket anti-AI across the board no matter what. I just don't think that providing them to newcomers who don't have a sense of our AI guidelines is a good idea. (For more on that reread GreenLipstickLesbian's comment)
Some things I don't think I've seen mentioned:
  • Right now, this and other task suggestions seem to target tagged/templated articles with high view counts. I think both parts of this are a mistake. People template articles because they want to make them better and, by and large, the result has been templated articles getting either worse, or impossible to tease apart good vs. bad edits due to the sheer volume of them. Something like this might look good from an edit metrics standpoint but from the standpoint of doing patrol it is a lot of patrol work -- and that's just the entry point to one article. Something like that can easily spawn 5 more tabs to check if someone was indeed doing high-volume AI edits. And of course the higher the view count, the higher stakes any mistakes become. (I will say that Paste Check has been somewhat helpful, it technically creates more patrol work but it's more like pointing to patrol work that would exist no matter what)
  • If the text parsing can't be fixed (Though it should be), the fail states are systematic enough that you can probably just regex filter somewhere in the process.
  • Controversial material needs to be blacklisted for reasons that should be clear. Admittedly the three examples above were somewhat cherry picked to demonstrate that point, but they were not especially hard to cherry pick. (The third one was just CTRL-F "partisan" because I already knew LLMs have issues with that, and sure enough there they were.) This should also be low-tech. Something like blacklisting CTOP and BLP articles -- I know Simple Summaries blacklisted BLP articles, though there were a lot of holes in that implementation -- and more importantly, an additional filter at the sentence level that is just blunt keyword-based stuff. This should also make MOS:GEO suggestions much safer since it's really hard bordering on impossible to predict every way that can go wrong.
  • I don't know what level of LLM familiarity the people working on this have, but I assume it is more in the LLM-coding-in-general realm, less the intersection of LLM output and Wikipedia. So, seconding getting in touch with the people at WP:AIT, they've been iterating on stuff like this for a while and have a sense of what does and does not work. I also can't speak for everyone, but while the people on AI Cleanup are less likely to be open to AI and even less likely to have any spare time to help out with stuff, they do have more of a front-row seat to AI edits and edit suggestions than basically anyone else. Most if not all of the issues above were predictable when you've seen a lot of LLM edits on Wikipedia. I have only found one common LLM edit problem that hasn't shown up in these (which is actually kind of interesting from a model standpoint). Feel free to email me, I have a great deal of collected data on this.
  • This is cheating since I've mentioned it, but "more QA" seems to be a constant across the board for all features. Another place where getting in touch with AI folks can help because spot checking goes a lot faster when you know what to look out for.
(Also, I know this isn't something you have any way of knowing about, but I prefer to not be pinged to ongoing discussions I am participating in, it just generates notifications and emails that pile up). Gnomingstuff (talk) 18:28, 19 August 2026 (UTC)Reply

Where can you see all of the model-generated MoS suggestions?

[edit]

Okay. Linked here you will find a spreadsheet that contains the batch of ~6,500 experimental LLM MoS suggestions as they appear in Suggestion Mode for people who enabled experimental suggestions.

Within the spreadsheet are two categories of suggestions: those related to simplifying language and MOS:GEO. We've removed all of the NPOV suggestions based on the feedback y'all have been helpfully sharing here.

How do these suggestions look to you? What are examples of specific suggestions that you find to be unhelpful, confusing, and/or just plain wrong? As you're going through these suggestions, what broader patterns are you noticing/conclusions are you reaching about these two categories of suggestions?

With this feedback in-hand, we're thinking we (staff and volunteers) can take a step back together and decide how and if we should move forward with this particular set of suggestions.

Of course, if you find any part(s) of the spreadsheet are unclear, please let us know.

A couple of notes:

  1. We welcome feedback in whatever form is most convenient for you. E.g. sharing directly in this discussion, enabling experimental suggestions in Suggestion Mode (see instructions) and offering feedback through the UI, etc.
  2. Some suggestions may be present in the spreadsheet and not visible when you edit the actual article with experimental suggestions enabled. This is because this spreadsheet contains suggestions that were generated in a batch offline. In the time since, some articles have been edited in ways that make the suggestions obsolete.

PPelberg (WMF) (talk) 22:35, 19 August 2026 (UTC)Reply

Thank you! Will take a look Gnomingstuff (talk) 22:43, 19 August 2026 (UTC)Reply
You bet and thank you! PPelberg (WMF) (talk) 22:49, 19 August 2026 (UTC)Reply
Is it worth implementing some sort of minimum length check for simplify language? I am not sure how much shorter "Scorer for Crystal Palace; Terry Fenwick" or "Town in Faisalabad District" can be. The model seems to feel parentheticals are a lot more complex than they are in reality: "Romelda Aiken-George (née Aiken}; born 19 November 1988) is a Jamaican netball player." It also appears to be picking up some template code or something ([ { "List of leaders of Markazi Jamiat Ahle Hadith": "Order" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "1" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "2" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" }, { "List of leaders of Markazi Jamiat Ahle Hadith": "" } ]) as well as picking up lists as prose ("See also
List of prime ministers of Pakistan List of presidents of Pakistan Chief Secretary Khyber Pakhtunkhwa List of chief ministers of Punjab List of chief ministers of Sindh List of chief ministers of Balochistan". It's even picked up the reference list for 1994_Andhra_Pradesh_Legislative_Assembly_election as in need of simplification.
If the list could be refined to remove the simply technically non-applicable, it would be easier to parse. I would be interested in a few examples worked through by editors here who have spent a lot of time looking into language simplification (@Femke?). For example, "It is the feminine form of the Late Latin name Clarus which meant "clear, bright, famous"" and "A space station (or orbital station) is a spacecraft which remains in orbit and hosts humans for extended periods of time" seem near fully concise to me, but I have not looked into this topic much. (Looking now as I copy these in, these may also represent examples of the model looking at too short a sentence and being confused by parentheticals.) CMD (talk) 02:29, 20 August 2026 (UTC)Reply
In this case, cross-referencing against the csv I have, the AI's "suggestion" is Fix the misplaced bracket and punctuation in the parenthetical expression for the maiden name. Which actually is a legitimate fix -- there's a stray curly brace -- but isn't related to simplification at all, see my comment about it being a catch-all.
I know it might be confusing and the spreadsheet here is supposed to mimic what editors see, but for QA purposes having those around without having to cross-reference might be helpful to pin down what's going on Gnomingstuff (talk) 03:04, 20 August 2026 (UTC)Reply
Thanks, same issue as a lot of the MOS:GEO examples then in the llm finding something that is out of its supposed scope leading to a confusing tag. CMD (talk) 03:07, 20 August 2026 (UTC)Reply
@Chipmunkdavis: thank you for reviewing! Responses to the points I see you raising...
Is it worth implementing some sort of minimum length check for simplify language?
Great question. Assuming we were to agree this suggestion was worth moving forward with, I think we could implement a way for you all to set a minimumCharacters value as is currently available for Reference Check. cc @DLynch (WMF) who can say definitively whether this is feasible.
The model seems to feel parentheticals are a lot more complex than they are in reality... as well as picking up lists as prose...
Great spots and noted. We'll investigate whether there are things we could do to exclude these two category of issues you're describing. PPelberg (WMF) (talk) 23:35, 20 August 2026 (UTC)Reply
Sure, there's two ways we could do it. The equivalent to the reference check would involve doing it client-side -- you'd set a config value in Editcheck-config.json on-wiki and we'd filter out suggestions that we get from the API that're below whatever length you specify. The other way would be to configure the model to not generate those suggestions in the first place, which would be more-ideal overall in terms of work saved. Harder for you to configure though. DLynch (WMF) (talk) 23:41, 20 August 2026 (UTC)Reply
By the way, for the "simplify language" suggestions I was wondering if existing small language models could do a similar job, so I bashed together something (I used the existing WMF readability model, m:Machine learning models/Production/Multilingual readability model card, though it isn't really designed for sentence level ARA... so possibly a way to improve things if a more suitable model can be found). I did have Gemini 3.6 Flash screen the about 2/3rds of the suggestions and fill out the description field (columns are same as the CSV instead of the spreadsheets) since I haven't investigated which features produced the scores, not sure how useful those descriptions would be. I've posted the suggestions it generated here if anyone is interested in looking at them and comparing to the ones generated by GPT-OSS. Alpha3031 (tc) 15:16, 22 August 2026 (UTC)Reply

What if communities don't want certain suggestions on their wiki?

[edit]

If a wiki does not want a certain suggestion to be shown to anyone on their wiki, any administrator or interface administrator can disable it directly, without requiring any code changes/action from WMF staff.

In practice, this would look like someone visiting MediaWiki:Editcheck-config.json, identifying the Edit Suggestion or Check they'd like to disable, and changing ShowAsCheck or showAsSuggestion from True to False. We're of course happy to help out/clarify any confusion. Although, the goal is for this system to be intuitive enough for you all to adjust it based on what you're seeing in practice and the expertise you've developed.

In addition to toggling Checks and Suggestions ON/OFF, there are a range of other ways you all can customize them:

  • Who sees them. account limits a Check or Suggestion to people who are logged-in or logged-out and minimumEditCount/maximumEditCount enables you to target them based on someone's editing experience.
  • Where within an article they appear. ignoreSections excludes Checks/Suggestions from appearing within specific section headings and includeSections does the inverse, restricting a Check/Suggestion to only appear within the sections you define.
  • What kinds of pages and content they run on. ignoreDisambiguationPages makes it so a Check/Suggestion does not appear on any disambiguation pages, ignoreQuotedContent prevents them from appearing on text inside of quotation marks or blockquotes, and inCategory/notInCategory and hasTemplate/lacksTemplate scopes Checks/Suggestions based on the categories and templates present/absent within an article. Note: there's a task for enabling configuration based on properties of an article's talk page too.
  • How sensitive an individual Check is. These vary by Check/Suggestion. For Reference Check, minimumCharacters enables you to define how much new text someone will need to have added for it to get activated. For inference-based Checks/Suggestions like link suggestions, predictionThreshold enables you to set confidence threshold the model must reach before anything is shown to someone.
  • Entirely new, locally-defined Checks. TextMatch lets you define entirely new Checks, locally using pattern matching, without any code changes from us.

At any point in time, you can visit Special:EditChecks to see the current set of Checks and Suggestions that are available and how they're configured. This page is updated automatically as code changes are merged. In the future, we'd like to add metrics so that you have even more visibility into how things are working and identify what might need adjustment.

Might there be ways you'd like to be able to configure Checks/Suggestions that we do not currently offer? Is there information that you'd like to see on Special:EditChecks that's not currently visible? More broadly, does how we're thinking about this line up with how you're thinking about it? We're open and eager to any and all feedback. It's important to us that you have the visibility into, and control over, this system to make it work for your wiki (and the same for volunteers at other wikis).

---

i. See the distinction between Edit Checks and Suggestions that Marshall posted above. PPelberg (WMF) (talk) 23:51, 19 August 2026 (UTC)Reply

What kind of communication with communities has there been on this project so far?

[edit]

In a June meeting hosted in the Wikimedia Discord, members of the Editing and Machine Learning Teams shared (discord link) that we were experimenting with LLM-generated MoS suggestions.

In that conversation, we mentioned that this initial experiment would not involve surfacing any actual edit suggestions to volunteers. Instead, we would show edit suggestions in an "experimental" state that invited the people who opted-into seeing them to offer feedback about the quality of the suggestions. Then after that feedback process, we (volunteers + staff) could decide whether any of them were worth actually showing to people via a controlled experiment.

In the time between that call and now, we shared progress updates on the MediaWiki project page and had planned to announce the existence of this work and invite feedback about it on-wiki this week or next.

I appreciate that the above still led us into a situation where many of y'all were caught off guard by this work and as a result, trying to put the pieces together in real-time. This is not ideal for y'all and it's not ideal for us!

In this thread we've seen folks like @Alpha3031, @Chaotic Enby, @SJ, and @Kowal2701 helpfully suggesting that work like this would be better off were we to be in touch about it early and often. This way, we can get on the same page about what ideas might be non-starters and for those that we do deem worthwhile to pursue, to align on when and how we'll evaluate them together.[1] [2] [3] [4]

This all leads me to wonder…

Let's imagine we (staff) have identified a new LLM-powered suggestion that we think would be worthwhile to experiment with. When would you all appreciate us checking in with you all about it? How much information/example/context do we need to prepare before it's useful for a large group to evaluate an idea? Where do you think would be the best place for us to post about this?

Asked as another way: if we could do this all over again, what would you imagine the "ideal" process to be? PPelberg (WMF) (talk) 23:08, 20 August 2026 (UTC)Reply

Not sure this is a communication issue, so much as a content and quality issue. If a feature is good then communication about it will be received more positively no matter what form it takes, whereas if a feature is not good then there is no way to announce it in a way that will be received well. Gnomingstuff (talk) 03:22, 21 August 2026 (UTC)Reply
To me, the ideal process would be asking the community first what LLM-powered suggestions they would be most interested in and would find acceptable. It might also avoid issues with setting up suggestions that may swerve into CTOPs and CTOP-adjacent spaces that you might not be aware of. There's plenty of low hanging fruit for LLM suggestions that I think would have a reasonable amount of support, but any time you're going to have an LLM suggest things, especially to new editors, that have resulted in arbitration cases and site bans you're probably looking at the wrong things. ScottishFinnishRadish (talk) 11:25, 21 August 2026 (UTC)Reply
Hello Peter, agreed with SFR + Gnomingstuff that this is about quality and collaboration, and how to start with easy cases when implementing any new interface or workflow. Sharing examples as they are produced, and working in public on essential parameters like what eval is used and the threshold for what gets presented, would make it easy for people to give targeted feedback early and often, saving everyone time.
* Give more attention to the choice of areas that will be covered. Have a page for suggestions, with definitions slightly more detailed than a one-sentence description. (what is the scope of each, defined how, tested against which guidelines)
* One step in the review checklist should tease out what might be easy or hard or surprising in each area.
* Spend time on better evals. Have human evals along with LLM-judge evals of initial suggestions.
* Use a tool for bulk elicitation and evaluation of suggestions (like an editable spreadsheet); working through the VE interface is too slow for meaningful human review at scale, and when reviewers find a type of error they will want to find many instances of it to fully characterize what's going wrong.
* Fine tune the models used on feedback.
This doesn't need to be a large group; small groups working in public on a predictable cadence is fine. The above is helpful regardless of what is generating the suggestions (LLM or other tool). It shouldn't be 'checking in' ⸺ this is part of the editorial workflow! There's no lack of interest in finding low-hanging fruit that works. I propose a better starting premise would be "we (all) have identified a suggestion that may be worthwhile to experiment with"... which is possible once there is an active page for suggestions.  SJ + 15:37, 21 August 2026 (UTC)Reply
Seconding all of these Gnomingstuff (talk) 06:14, 22 August 2026 (UTC)Reply
Also, you might let people test + give feedback on the interface separately from particular suggestions. For instance, using this to highlight {cn} instances on a page or other inline flags that are already present in the wikitext but harder to discover. That could happen continuously while studying low-hanging suggestions above.  SJ + 15:37, 21 August 2026 (UTC)Reply
@ScottishFinnishRadish and @Sj: thank you for offering these clear recommendations for how we might work more effectively together going forward.[i] And thank you Gnomingstuff for stating clearly that no amount of communication substitutes for the underlying quality of the work.
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion. Until then, thank you for the thought and attention you've offered this week...we continue to learn a great deal in the process!
---
i. E.g. create a local project page where we can generate and evaluate ideas for new suggestions together, make evaluation easier to do at-scale, avoid contentious topics when considering suggestions to pursue, etc. PPelberg (WMF) (talk) 23:25, 21 August 2026 (UTC)Reply
Next week, you can expect me to follow up with the concrete ways we're thinking about integrating the feedback people have shared in this discussion.
Hi y'all – there are at least two ways the work will proceed from here:
First, on the current batch of MoS Suggestions: we’re going back through this discussion to compile a list of issues that we can investigate and hopefully, use to improve the MOS:GEO and Simplify Language suggestions (spreadsheet).
Second, on process: we're going to draft a proposal for where/how we can evaluate ideas for new suggestions together (e.g. copyediting) and metrics we can use to assess the quality of suggestions once a dataset is available.
You can expect another message to Village Pump as soon as we have updates on the above.
In the meantime, thank you all for this helpful feedback. PPelberg (WMF) (talk) 18:25, 26 August 2026 (UTC)Reply
 SJ + 13:22, 27 August 2026 (UTC)Reply
Re: communication, I think it depends on the context. Is it something suggested by several editors on-wiki? Is it very similar to an existing feature? Is it similar to an existing user script? Go for it. Have similar ideas been controversial in the past? Is it a completely new approach to something? More discussion with the community as you design and develop it would probably be good. If in doubt, just drop a "hey, what do you guys think about an AI tool to detect NPOV issues?" on a noticeboard somewhere. It doesn't need to be some elaborate presentation. (I know big communities and corporate cultures can both make some of those things hard, but aspirationally I wish we had a culture—both here and at the WMF—that made that not a big deal and encouraged the communication.)
The best place is on wiki, either on this page or (for more specific features) on a relevant talk page. This is where we all are, and the same can't really be said for anywhere else. (Disclaimer that I can't speak for other projects.) And if you're not sure, it's generally pretty easy to ask "where is a good place for this" or just to drop a link on a few pages. LittlePuppers (talk) 03:50, 22 August 2026 (UTC)Reply
I'd advise against relying on suggested by several editors on-wiki. I'm sure you could find more than several who would support any and all LLM implementation across the board. "Consensus between multiple experienced editors" might be better phrasing, but I think similarly to existing tools that have general acceptance is a safer metric for proceeding without additional discussion. Anything that's novel should probably receive community input. ChompyTheGogoat (talk) 21:09, 27 August 2026 (UTC)Reply

Possible RfC

[edit]

Given the WMF seems inclined to continue developing these tools regardless of what we say, I propose we hold an RfC establishing that LLM tools may not be deployed without an explicit consensus from the community. If editors think this may be a good idea, I suggest the following as the first draft of the question:

Should we require an explicit consensus before the WMF is permitted to deploy tools involving the use of an LLM? The WMF may deploy tools for user testing, so long as all of the following criteria are met:

A: Testing of the tool concludes after no more than 12 months, after which the tool must be removed unless there is a community consensus for either extended testing or full deployment.
B: Use of the tool is restricted to editors who have opted-in
C: Use of the tool is restricted to editors who have been granted the LLM-tool user right. This would be a new user right created and granted by admins at WP:PERM.

Does anyone see issues with this proposed question? Are there any revisions? BilledMammal (talk) 07:42, 21 August 2026 (UTC)Reply

I'm not fundamentally opposed to this, but I think we need a lot more clarity about what this new LLM-tool user right would mean. A naive reading of the name would be "This user is exempt from WP:LLM" which I suspect is a broader interpretation than you intended. RoySmith (talk) 12:53, 21 August 2026 (UTC)Reply
This seems like policy creep, and not a good approach (we shouldn't have blanket limitations based on the "type of tool" involved). I believe it is also misdirected: the example above is testing what would be an entirely community-configurable tool. Opt-in is appropriate for things that are so new, and I'd say anything that messes with your margins should be configurable. But proliferating user rights is an anti-pattern, and should only be used as a last resort.  SJ + 16:21, 21 August 2026 (UTC)Reply
yeah, I think this may be premature for now, but worth revisiting if there are any disasters Kowal2701 (talk, contribs) 18:24, 22 August 2026 (UTC)Reply
Consistent with the current wording of NOLLM and with active work being done - so far successfully - at helping limit abuse, I would suggest some cleaving of administrative actions and content actions. Best, Barkeep49 (talk) 17:05, 21 August 2026 (UTC)Reply
Solely regarding access control (I'm undecided regarding the overall proposal): I agree with the other commenters that that creating a new user right isn't the best fit. I think tool-specific JSON lists of approved users that are fully protected could suffice, and be more adaptable to allow for per-tool authorization. isaacl (talk) 17:15, 21 August 2026 (UTC)Reply
The whitelist system works well for AWB/JWB. Certes (talk) 21:13, 27 August 2026 (UTC)Reply
I also think that the community should be able to control and configure the tools used in en wiki but I'm not sure an rfc is needed atm and its wording unfortunately isn't clear.
What is a "tool"? What else, other than edit suggestions, is in the scope?
12 months deadline is arbitrary: it's too long for a tool that the community actively opposes and may be too short for an iterated testing of a complex idea.
Some abuse/vandalism prevention tools by definition cannot be opt in. Alaexis¿question? 20:18, 21 August 2026 (UTC)Reply
It's hard to guess what is a tool, because the AI is developed secretly and added to Wikipedia's software or configuration as a fait accompli. Certes (talk) 21:13, 27 August 2026 (UTC)Reply
I agree with most of the comments above that this is not a well-defined RfC. In addition, the WMF have said this will be community-configurable if it makes it into production so it appears to be unnecessary anyway. Mike Christie (talk - contribs - library) 18:34, 22 August 2026 (UTC)Reply
I assume this is intended to apply to any future implementation that would integrate LLM features in any way, not just this one concept, which I fully support. Pushing it on us without community consent is inappropriate and exactly the type of forced AI we're seeing on every other platform. Wikipedia is supposed to be different.
I'd agree on amending the deadline - I'm not sure what an appropriate initial test phase would look like, but I'd recommend a limit aligned with that and only extended or fully implemented via consensus. Just enough to get a feel for it, not a full beta phase to polish it for release. People with more programming experience than me would probably have a better notion of the timeline.
I do think we need some limitation on access (especially if it is fully implemented) and I think Isaacl's suggestion sounds like a better way to handle it. ChompyTheGogoat (talk) 20:58, 27 August 2026 (UTC)Reply
I think that some level of access control is warranted, but the idea of getting individual approval for every tool that the WMF wants to test would impede testing numbers without much safety gain over a single user list for all beta tools.

Having a time limit on testing is strange, but my exact opinion on it depends on how community consensus is defined here. I think something that both mitigates the issue i think you're trying to solve while not dictating the WMF's development calendar would be something like "Any tool in testing shall be disabled if a consensus is formed at the tool's thread on VPWMF that the tool is causing disruption. At which point the tool will only be re-enabled with consensus." Although, if we got anything near consensus that a tool in testing is disruptive, I think the team working on the tool would disable/fix it very quickly. These are people who chose to work on mediawiki here, not WMF management. MetalBreaksAndBends (One for all) 01:32, 28 August 2026 (UTC)Reply
I think it's only suggesting permissions for LLM tools that would potentially be more prone to abuse, not for any and all in testing. ChompyTheGogoat (talk) 07:13, 28 August 2026 (UTC)Reply
There isn't any limiting language in the comment, so I (and most people, probably) would think it would apply to all LLM tools. MetalBreaksAndBends (One for all) 23:42, 28 August 2026 (UTC)Reply
That is what I meant - all LLM, not all tools period. If we end up with so many LLM tools being tested that they're hard to keep track of I'd consider that a problem unto itself. ChompyTheGogoat (talk) 02:50, 29 August 2026 (UTC)Reply
I'm not saying the issue is that they could be hard to keep track of (though over a long enough time period it could), I'm saying that applying for every beta is an unnecessary hassle. MetalBreaksAndBends (One for all) 03:38, 29 August 2026 (UTC)Reply
I think I would share the view that a full, formal RFC would be unnecessary if prior discussion shows clear consensus to implement, for example (assuming it's advertised to a noticeboard like this one or the cleanup wikiproject). I would expect any community members participating in such discussions to be sufficiently in touch with current attitudes towards LLM tools to bring up anything that would seem potentially controversial and standard pre-RFC discussion processes should function fine (w formal closures and move to full RFC if no clear consensus or if there is consensus there should be an RFC in the specific discussion). Alpha3031 (tc) 04:05, 29 August 2026 (UTC)Reply
Do you mean a full rfc for each tool or an RFC for this proposal? MetalBreaksAndBends (One for all) 06:25, 29 August 2026 (UTC)Reply

Hey WMF: nice job on communicating in this section

[edit]

I think this conversation has been healthier than some other recent enwiki–WMF conversations. Just wanted to say thanks to the WMFers that are participating here and doing things like compromising (i.e. getting rid of the NPOV model), replying a lot instead of making one polished statement then leaving, and speaking clearly and honestly.

It's tough because on some issues, WMF and enwiki are very out of sync. So even if everyone does everything right, these conversations may still be tough. But doing things like compromising, having conversations with us that aren't just statements, and speaking clearly and honestly are definitely a good approach. Please keep it up. –Novem Linguae (talk) 19:23, 22 August 2026 (UTC)Reply

Absolutely seconding this! Really happy to see WMF folks take community feedback into account. I understand the task can be much harder than it seems at first, especially as these discussions get sprawling and it can be hard to find a thread that unites the whole range of community opinions together, let alone incorporate it in the team's plans for the project.
The way you managed to navigate it was a very positive surprise, and I'm looking forward to more productive exchanges from both sides! Chaotic Enby (in solidarity · talk · contribs) 19:45, 22 August 2026 (UTC)Reply
Yes, absolutely. Ymblanter (talk) 08:10, 23 August 2026 (UTC)Reply
Absolutely agreed on all of this Gnomingstuff (talk) 19:01, 23 August 2026 (UTC)Reply
+1. Can't speak for anyone else, but pretty much all my complaints about WMF's poor communication are limited to the Trustees and a few Officers. Everyone else at the WMF seems to communicate fine. MMiller and PPelberg are two usernames (among others) I've grown very accustomed to seeing regularly on-wiki, they've communicated often and effectively with volunteers for years. Levivich (talk) 20:22, 23 August 2026 (UTC)Reply
Thank you -- we're glad to hear it -- we are trying hard! Although there are going to keep being times that we all disagree on ideas, the most important thing is that we can discuss constructively to figure out how to best improve/adapt the wikis. Thank you all for being here for these conversations, on top of doing all your usual wiki work. MMiller (WMF) (talk) 05:20, 24 August 2026 (UTC)Reply
I am of the opinion that most -- maybe all -- WMF employees who are not in top management are competent, helpful, want to do the right thing, and are eager to communicate. I attribute the stonewalling we often see with a perfectly reasonable fear that actually having a dialog with the volunteers will never help your career and just might get you fired. If only there was some way that WMF workers could organize and join an entity that works to protect them from being unjustly fired... Let me know if anyone has ever heard about something like that. --Guy Macon (talk) 12:38, 24 August 2026 (UTC)Reply
I find this post amusing. But as a Wikipedian, I can't help but be myself and note that if you define top management as something beyond "C Suite" Marshall is would probably be considered top management. Best, Barkeep49 (talk) 14:47, 24 August 2026 (UTC)Reply
+Whatever we’re on In solidarity Wikipedian12512 (Talking is fine | contribs) 01:05, 6 September 2026 (UTC)Reply

Comment from trustee

[edit]

Hi everyone. Awhile ago, I sent messages to individual trustees. As far as I'm aware, only one person has responded, but I found one of these comments somewhat insightful. They've said that they're aware of these discussions (good), which likely means other board members are too. So I think the inaction is a choice. Especially after reading this:

Hi, It's not by accident that those board members are responding, it's by design. The other problem is that it becomes very difficult to solve anything when people are more focused on winning an argument than actually solving the problem. Did you see the message above? The user doesn’t even greet, they come in guns blazing, already prepared for a fight. How do you engage cogently with that level of anger? Because even if you explain that we are not union-busting and provide the reasons and back it up with evidence, the response will simply be: “You’re lying, you are union busting. Then they present their own reasons, sometimes mixed with misinformation and then you have to counter that by proving that part of what they are saying is actually not true, and we end up going in circles. At the end of it all, no problem has been solved. You may have won the argument, but the real question is: what did winning the argument actually solve? So me and the other board members repeating the same things that the 3 board members have already said just adds to the noise than solving a problem really.

While that's an incredibly frustrating response, it is at least a substantive one. It feels like it was written by a person and not a corporation. Maybe someone who is not me will be able to convince the board that this is not how you solve this. Clovermoss🍀 (talk) 01:28, 20 August 2026 (UTC)Reply

Don't read too much into that last sentence. It was a general statement, not a suggestion to have a bunch of people en masse go to that guy's talk page. That hasn't happened yet, I just wanted to make that extra clear upon a reread of what I said, especially since Meta seems to have some arbitrary unwritten rules about that that might get people blocked (m:Universal Code of Conduct/Coordinating Committee/Cases/2026/Chilling effect on Metawiki). I don't want people to get in trouble and I wouldn't feel comfortable knowing this and just not warning people about potential risks in that capacity. I'm assuming people will have opinions about this, but I'd rather the bulk of that discussion take place here, rather than making someone who finally said something feel cornered. Clovermoss🍀 (talk) 03:04, 20 August 2026 (UTC)Reply
So me and the other board members repeating the same things that the 3 board members have already said
if only there was something they could do instead of repeating themselves! ah well Gnomingstuff (talk) 03:07, 20 August 2026 (UTC)Reply
To be fair, some people do "come in guns blazing, already prepared for a fight" and in those cases I completely understand not responding.
However, I and other users have asked a lot of questions in a lot of ways over the years, and the result has (with a few exceptions, almost always related to either a huge discussion elsewhere, a petition, or an article in The New York Times) has been the exact same silence. Rarely, you get a non-answer in corporate speak, followed by silence when you try to have an actual conversation with them.
Here is what has been tried:
  • Be a total jerk and start off with an insult: Result: no response.
  • Be super polite and deferential: Result: no response.
  • Ask one person at the WMF, once: Result: no response.
  • As above, but ask again every month or two for years: Result: no response.
  • As above, but ask all sorts of people at the WMF: Result: no response.
  • Ask on multiple pages on multiple projects: Result: no response.
  • Have the person asking be a long term contributor who has never said a word about the WMF before: Result: no response.
  • Have other editors respond to any of the above with "I would like an answer as well" comments: Result: no response.
  • Ask for a response by email with a promise never to reveal that they talked to you: Result: no response.
There are two exceptions that commonly occur. Jimbo often responds. And a boatload of ordinary Wikipedia volunteer editors often post answers, some of which are quite helpful.
I invite the attentive reader to consider what the common factor in all of the above is, and how it relates to any "you didn't get an answer because of the way you asked" or "you didn't get an answer because of who you are" claims. --Guy Macon (talk) 06:43, 20 August 2026 (UTC)Reply
"Jimbo often responds" but rarely says anything useful or believable. Like others said on his talk page recently, " I think there's a general trend of miscommunication which is mostly on you. You have a general attitude of dismissing opinions that you don't agree with, instead of engaging with them." or (from another commenter) "If you don't care to substantively engage with my comments, that's your prerogative, I suppose, though it's frustrating given that your refusal to address most of them comes in a message where you continue to insist you don't have a general attitude of dismissing opinions that I don't agree with, instead of engaging with them. " Fram (talk) 08:00, 20 August 2026 (UTC)Reply
Just look at his flip-flopping about the date Littler Mendelson was hired.
  • 9 August: "As of today I do not know exactly when they were hired. Obviously I am unhappy about every aspect of that."
  • 9 August: "I can't promise anything of course, and it isn't for me to decide if a date like that is released publicly. But my strong recommendation is for maximum transparency possible and at least at this moment I can't think of any reason why that'd be something to keep private."
  • 10 August: "I have looked into it and spoken further with the Foundation leadership team to try to learn more about where things stand. I am satisfied with the specific work the Foundation has asked Littler to do. " bu nothing about that date
  • 10 August: "I haven't suggested anything about when Littler was hired - I actually don't know when they were hired and so I've avoided saying or suggesting anything about it at all. For all I know it could have been January or it could have been much later. I just don't know and I also am totally unclear on why it's important - although I will try to find out." (it's become unimportant in one day, and clearly wasn't asked at the 10 August meeting then?)
  • 10 August "The point is, why would I have that information, at least in general. When I have my next meeting with WMF I'll ask if they can release it. But maybe you can let me know why it's important."
  • 11 August: " I don't yet know exactly when Littler was hired, "
  • 13 August: "Yes, that's at the top of my list when I meet with the WMF. " (about getting to know the date they were first hired)
  • 13 August: "A meeting is being scheduled with the C team and the board, and at that meeting I'll get the best opporunity to pass along the questions that have been asked, questions about the things that I don't know yet."
  • 15 August "I'm with you." (about the importance of establishing a timeline)
When he needs to calm things down, he agrees that it is important and will find out. When he do has the chance to find out, he doesn't and reappears saying that he doesn't understand why it would be important. After pushback, he again finds it important and puts it at the top of his list. And then again nothing... And this is just one obvious, concrete example. But his talk page is full of unanswered, deflected, ignored questions. Fram (talk) 08:30, 20 August 2026 (UTC)Reply
My saying that Jimbo responds more often than the rest of the WMF combined is sort of like saying that, of the Three Stooges, Curly is the intellectual stooge. It isn't a high bar.
Related: Lies, Damned Lies, and Union-Busting Statements -- Medium
Pay careful attention to the section titled The Lie: “There are no union-avoidance law firms involved”. --Guy Macon (talk) 09:17, 20 August 2026 (UTC)Reply
Well, we would have thought... no reply so far, and as he will be quiet until 1 September, and the WMF is also quiet until after the election is finished, we just happen to not get an answer on this very simple question, weeks after he was "very unhappy" and recommended "maximum transparency" and so on. Anyone surprised? Am I too cynical if I think afer the election results are in, comments from the board or the WMF will be that it is "time to move on" and we need to "pull together and move forward as a community"? Fram (talk) 16:34, 24 August 2026 (UTC)Reply
They're gonna say they can't talk about anything due to the confidentiality of ongoing CBA negotiations. Anyway, what is there to talk about? The c-suite and board have made themselves exceedingly clear in response to volunteer inquiries, IMO. I'm not sure there is anything to do other than have new trustee elections ASAP so the new trustees can bring us a new c-suite. Levivich (talk) 16:42, 24 August 2026 (UTC)Reply
Wow, and that comes from a community appointed trustee... Ita140188 (talk) 09:54, 20 August 2026 (UTC)Reply
With a background in communications ironically ~ In solidarity 🦝 Shushugah (talk) 17:08, 21 August 2026 (UTC)Reply
Either of the candidates who were removed from the 2025 WMF Board of Trustees election with two days' notice (see community feedback) would have done a better job of communicating than this Board member has done here. — Newslinger talk 07:40, 21 August 2026 (UTC)Reply
Mark my words: with this new "Board qualifications" effort, we will never again have the opportunity to elect a non-pre-approved candidate. They are shutting the doors behind them. Levivich (talk) 21:06, 21 August 2026 (UTC)Reply
  • Given the fact that the WMF now only allows candidates that they approve, I think we should post a petition asking them to add "none of the above are acceptable" to the ballot. Ideally, if NOTA wins, that should trigger a new election with the unacceptable candidates excluded. Eventually they will run out of handpicked candidates willing to rubber stamp what the WMF has already decided to do rather than respresenting the community. If they wont allow NOTA on the ballot, it will be time to boycott voting. Do they currently say how many people voted? If so, suddenly deciding to keep that information secret after a voter boycott gets organized will be quite suspicious. --Guy Macon (talk) 17:25, 24 August 2026 (UTC)Reply
    Good idea. I already refuse to vote in elections which are limited to WMF-approved candidates or are for unwanted committees which should not exist. Certes (talk) 18:16, 24 August 2026 (UTC)Reply
    I don't remember how the securepoll is set up -- whether you have to vote "yes" for a certain number of slots, or if you can vote "oppose" to all candidates (which, IIRC, is possible for, e.g., arbcom elections; not sure if it's the same for trustee elections).
    One thing I urge the community to do for the next trustee elections, whenever they are, is require certain explicit pledges from the candidates, and refuse to vote for any candidate (whether pre-approved or not) who does not agree to make the pledges asked of by the community.
    I'm not sure exactly what pledges would have community consensus, but something like: require the Board liaison committee to answer all inquiries on the meta Board Noticeboard within something reasonable like 7 days. I would add some other things like: pledge to rescind the prior Board resolutions that allow the Board to pre-select trustee candidates; pledge to reform the Board Code of Conduct so it doesn't prohibit the Board from effectively exercising oversight or communicating with the community; pledge not to hire externally a CEO who doesn't have prior CEO experience (it's mind-blowing to me that the Board picked as CEO of a hundred-million-dollar, hundreds-of-employees, int'l non-profit, someone who has never been CEO of any similar or even smaller-sized organization ... an org of this size should be nobody's first run as CEO, unless they're promoting from within); maybe pledge to pass a resolution requiring all vendors (e.g., law firms, accounting firms) to be mission-aligned (that's kind of fuzzy to determine/enforce, tho)... whatever has consensus, we should figure out what actions we want our next Board to take, and then only vote for candidates who pledge to do it. Levivich (talk) 20:16, 24 August 2026 (UTC)Reply
    I think the problem is that the WMF allows to vote only for (arbitrarily) approved candidates. By definition then, these are not independent candidates and can never represent the community. Essentially, voting becomes an empty exercise. We have seen the result with the current trustees, that either refused to engage at all, or reply like this Ita140188 (talk) 06:33, 25 August 2026 (UTC)Reply

What I was told when I was on the board was that they did not want us to communicate publicly with the community as they were concerned the community might misinterpret our personal position as being the position of the board. This was similar to why they did not want us speaking with staff in a personal capacity either. I however consider most of you, aswell as most staff bright enough to know that an individual trustee does not speak for the board (unless they say otherwise). And I imagine there was also concerns about us revealing "private / confidential" details. Jimmy is simply given more leeway as the rest of the folks on the board / execs are unlikely to try to remove him for a claimed breech. Doc James (talk · contribs · email) 03:46, 29 August 2026 (UTC)Reply

One month after his first comment on the date Littler was hired, we still don't have the answer to that very simple question. But rest assured, "I expect something from the WMF to be posted here soon on that question. I have encouraged them to be as transparent and open as possible, to the point of it feeling uncomfortable, and they don't disagree! ". Meanwhile, the Board postedthis, stating "The Board will continue to provide oversight and long-term direction to Foundation leadership as this work proceeds." and other similar corpspeak. No replies were given to any questions posted there, and the Foundation Board noticeboard is an absolute wasteland. Checking to see which board members actually read that board gave one reply in 19 days. Fram (talk) 08:11, 9 September 2026 (UTC)Reply

Technically board members are supposed to be able to get and see any document they want. In reality, unfortunately the organization has tried in the past to withhold details despite being directly asked by a board member, see Knowledge Engine (search engine). I am not sure if that is also the case here. But there is definitely the possibility that an individual board member has asked and been stonewalled / told they need to go through some process. Doc James (talk · contribs · email) 14:24, 9 September 2026 (UTC)Reply
Outrageous. In California (where WMF is headquartered), Florida (where WMF, Inc. is incorporated), and I think most if not all of the rest of the U.S., Board members have a right to inspect the corporation's books and records (whether for-profit or non-profit). Otherwise, they'd be unable to carry out their fiduciary duties of managing the corporation. Levivich (talk) 16:41, 9 September 2026 (UTC)Reply
My impression from talking to board members is they often think that doing what they need to meet their fiduciary duties would be a breach of their fidicuary duties. Maybe I'm the misinformed one, but you're one of the handful of people who I've talked who have experience with boards outside the WMF, and everyone I talk to has felt similarly outraged when I point to certain things being the norm. Clovermoss🍀 (talk) 16:45, 9 September 2026 (UTC)Reply
These sample nonprofit bylaws from Stanford have a section called "Inspection by Directors" that says "Every director shall have the right at any reasonable time to inspect [the nonprofit's] books, records, documents, and physical properties." Here are four more sample nonprofit bylaws I found via a quick Google search; each have the same or a similar section: . The WMF's bylaws, however, do not have such a section. The right to inspect law still applies, though. But not putting it into the bylaws is... as the kids say, a choice. Levivich (talk) 19:15, 9 September 2026 (UTC)Reply
Sounds like something that should be brought up at the board noticeboard. You'd probably do a better job of it than me. Clovermoss🍀 (talk) 19:24, 9 September 2026 (UTC)Reply
No offense, but I'm not going to waste my time tilting at windmills, particularly given the extremely frosty reception I got the last time I tried to talk with the trustees (response 1, response 2, the others didn't respond at all). Others should do whatever they think is best, but for my part, I'm done trying to reason with this Board. Somebody ping me when, or if, they hold the next trustee elections. The pending issues we've discussed: reforming the Board code of ethics, the bylaws, the election-vetting process... are better off being discussed with the next batch of trustee candidates than with the current trustees IMO. Levivich (talk) 20:02, 9 September 2026 (UTC)Reply
Fair enough, I just hope for a future where people don't get so jaded they don't feel like it's wasted effort to even try. But it's on other people to restore that trust. Clovermoss🍀 (talk) 16:28, 10 September 2026 (UTC)Reply

Wishlist 2027, invitation to share feedback

[edit]

Hi, I’m Sonja and I lead some of the teams at the Foundation who will be responsible for picking up wish work under the new wishlist process. As you may know, the Community Wishlist started out as an annual process through which Wikimedia contributors submit and vote on technical improvements they would like the Wikimedia Foundation to work on. The main goal of it is and has been to improve the editing experience by making changes and features the community asks for specifically.

In recent years, the process behind the wishlist has changed, and we’ve heard from many of you that it no longer meets many community members’ needs. So now, the Foundation is designing a new process with the community to improve how wishes are triaged, voted on, and prioritized in a way that is transparent, balanced across project families and language editions, and takes into account what the Foundation can deliver.

I would like to get community input specifically on these three stages of the Wishlist process:

  • The triage stage, meaning how wishes are fleshed out, organized and filtered prior to voting
    • We recommend to have a working group, including volunteers from various wikis and Wikimedia Foundation staff to work through this together
  • The voting stage, including who may vote and how votes are structured
  • The post-vote stage, including how to bring equity into what work is prioritized
    • One way to do this is to rank wishes within 3 categories: Large Wikipedias or covering all wikis, small and medium-sized Wikipedias, and sister projects, so that top-voted wishes from smaller projects also get attention

This message is an abstract of the full proposed process. As you read the proposed ideas on Meta, please speak up about whether you think this is working well or if there are ways to make it stronger. Regarding the timeline, this consultation is open for two weeks. You can post your feedback on Meta or in response below.

For this year’s cycle, we plan to have the wish submission period in late October/early November and the triage process completed by late November. To respect the end-of-year holiday season, voting would happen in early to mid January. This first voting cycle is meant as a first step to try out a new process, and there will be more opportunities to provide feedback along the way, so that we can figure out the best process for future years together. SPerry-WMF (talk) 17:41, 27 August 2026 (UTC)Reply

Is there an overview that contains a list of all the things that made the wishlists and whether they were implemented? Ideally, such a list would break down whether an unimplemented wish is [A] something the WMF still wants to do but hasn't done, [B] something that is possible but the WMF decided not to do it, and [C] things that the community wished for that are impossible.
That is assuming, of course, that the WMF can identify when something is impossible. Remember when the Flow team promised us "No edit conflicts, ever" and I pointed out that this was formally proven to be impossible? (See Wikipedia talk:Flow/Archive 4#No edit conflicts? and Wikipedia talk:Flow/Archive 4#Brewer’s Conjecture.) Good times... --Guy Macon (talk) 17:08, 28 August 2026 (UTC)Reply
Results from previous surveys can be found on the respective survey results page. We currently don't have a running document of all wishes and those explanations, but we can keep this suggestion in mind moving forward. One problem with the full summary you're suggesting is that status labels have changed over the years, and with them decline reasons. We're trying to create a clearer set of statuses and are actively discussing how declining of wishes should be handled. Do you have any thoughts on that? SPerry-WMF (talk) 22:41, 28 August 2026 (UTC)Reply
I have a specific comment and some general comments.
Specific to what you are working on, It seems to me that someone sitting down for an afternoon could translate the rejection reasons or at least comment on those old wishes so as to make the history helpful. What I am thinking is that when you efficiently solve a problem it goes away and nobody talks about it, but when something doesn't get immediately solved it remains an annoyance. This gives a false impression about how effective the team is by making a lot of the good stuff invisible. The list I described gives both equal visibility.
My general comment is about the nature of wishlists, based upon decades of addressing similar issues in industry.
Consider two wishes. They both seem to be roughly equal to the people making them but actually have wildly different difficulty. See [ https://xkcd.com/1425/ ] as an example. Maybe wish A is 5% more popular than wish B but takes a thousand times more work to solve. You need to figure out how to do the less popular thing that takes a few hours first, and somehow communicate to those making the wishes that some things that look easy are hard and some things that look hard are easy.
Or consider an example I gave before: The Abcom word counting template doesn't actually count words correctly. This has no effect on most users but the users it hits get hit hard while in the middle of an already stressful situation. The Arbs and clerks can't fix this -- they aren't developers and don't have the skills -- so they bodge up some crappy workarounds like saying "the clerks will eyeball the page and ignore the bogus word count". Your team could fix this in an hour or two. The most junior developer is able to do a reasonable job of counting words. But it will never, ever, get to the top on any community wishlist because most people never see the bug.
You should spend a significant percentage of your resources fixing these small, easy to fix problems that have been annoying users for years instead of spending most of your effort on big, sexy. popular, and exciting things and leaving a huge mountain of quality-of-life technical debt in the hands of unpaid volunteers. --Guy Macon (talk) 15:14, 29 August 2026 (UTC)Reply
When deciding what to work on next, there's a lot of factors. One is certainly how hard it is. But another is how many people the issue affects. Another is how much pain does the problem cause. I see word-counting arbcom statements as being pretty low on both the "how many people" and "how painful" scales and thus a perfect project for somebody to solve with a user script.
There's also the "how much collateral benefit will this bring?" scale. When doing work planning, it's not uncommon to look at something and say, "The user-facing benefit of doing this isn't huge by itself, but the work involved to do that will also result in refactoring this other gnarly thing which is a blocker for five other projects in our backlog, so it's worth doing". We, as users, typically have very little visibility into that aspect.
There's also the question of who's familiar with the code. Sometimes you go around the room and somebody says, "I'm all over that part of the system and know exactly what has to happen to implement this so I can knock it off in an afternoon". Sometimes you get a bunch of blanks stares and people muttering, "I didn't even know that existed; it'll take me a couple of days of exploration before I can venture an estimate of how much work it'll be".
Guy, looking at your userboxes, I would expect nothing I've said here will come as a surprise to you. It's cool that you know C and assembler and Forth. Me too, on all counts. RoySmith (talk) 15:43, 29 August 2026 (UTC)Reply
Thanks for your work on this. I'm glad there's a plan to have an iteration of the wishlist basically this year.
For this year’s cycle, we plan to have the wish submission period in late October/early November and the triage process completed by late November. To respect the end-of-year holiday season, voting would happen in early to mid January. I believe older wishlists allowed the creation of wishes and then voting on the wishes simultaneously. That is, in old wishlists you could create the wish during the voting period.
For this iteration of the wishlist, it sounds a bit like you are proposing a system where we can only create wishes before November, then a committee pontentially vetoes them, then we have to wait 2 months before we can vote on them? If this is the plan, then I think adding so much time to the cycle is not a good idea. The annual cadence of old wishlists got people to focus on the wishlist for a couple weeks -- this new system would require folks to focus on it for a couple months (create a wish at the proper time, wait 2 months, then market the wish so that it gets votes). I think it is really important to shorten the cycle in order to keep things nimble and to avoid bureaucracy. Perhaps all triaging should occur during and after the voting is completed, so that folks can easily create wishes during the voting period. –Novem Linguae (talk) 08:32, 29 August 2026 (UTC)Reply
believe older wishlists allowed the creation of wishes and then voting on the wishes simultaneously. That is, in old wishlists you could create the wish during the voting period. i dont remember this. We had vote creation combined with the community feedback phase, and some people would vote before the vote had started, because for many people it was difficult to understand what they were supposed to do. —TheDJ (talkcontribs) 17:49, 29 August 2026 (UTC)Reply
Yes, historically it was intended that there'd be a "filing and discussing wishes" phase and then a "discussing and voting on wishes" phase. Early voting happened but was various levels of discouraged. AntiCompositeNumber (they/them) (talk) 19:21, 29 August 2026 (UTC)Reply
This is all really helpful, thank you for weighing in.
The timeline we picked was meant like this: wishes can be submitted anytime between now and early November, but we'd run banners for the last 2 weeks of the submission window to get people to participate. Then we'd do triage for 2 weeks on all the wishes, meaning the working group would look through wishes, size them, ask anyone who is following the wish questions and discuss and document tradeoffs, and in some cases where it would be very clear that the wish would not be possible to be picked up (see potential reasons on under Triage activities and discussion with new ideas and perspectives on the Talk page), the wish might be declined. Then we'd have a voting period, widely advertised with banners, in January. After that we'd review the vote count to see if anything should be adjusted, for example to bring in a top voted wish from a smaller project, and we'd share the list we'd plan to work on with the community, again with an invitation for feedback (open for a week or two), before we start implemetation.
The reason for spacing it out like that is merely logistics: closing wish submission makes it much easier to triage wishes, because you don't have an ever growing list, and having the voting period in January was simply to respect the end of year holiday season. We're also a bit in a time crunch this time around, because we want to have a prioritized list by early February at the latest, so that we can start including them in our roadmaps for the current fiscal year (meaning between Feb and June). That part will be different in future wish years, because wishes voted on in January would not get worked on until the start of the next fiscal year in July. Aligning the vote with our annual planning process ensures that we free up the necessary resources to deliver on wishes alongside other work we have planned. This year is different, because we want to action wishes as quickly as possible under the new process. Aside from picking up wishes pretty much immediately after the vote, we will also bring them to the table when we start our planning process for the 27/28 fiscal year in February, so this round will feed wishes into this and next fiscal year.
With regards to not closing wish submission until we close the voting period: that makes sense, and you make a good argument to keep the periods closer together. Let me think that through a bit and come back to you with an alternate timeline early next week. SPerry-WMF (talk) 23:00, 29 August 2026 (UTC)Reply
Thanks for the detailed response and for taking the feedback onboard.
If you keep the phase system (submitting, then triaging, then voting), it could make sense to schedule it so it doesn't intersect with the December holiday season. Then less of a break would be needed in the middle, and the gap between submitting and voting could be shortened.
I think you all are envisioning the wishlist being independent of the annual plan. If that turns out to be true, it can be scheduled anytime. But assuming that it does need to be scheduled before annual planning season, perhaps Oct-Nov (before December) would be a good time period to shift it to. –Novem Linguae (talk) 10:13, 30 August 2026 (UTC)Reply
Those are good points – I think for future years it would make a lot of sense to move the submission and voting period to January and February, which would help us keep the timing between those periods tighter and it would reduce the time between vote and wish implementation as well. We specifically want to bring the results of the vote to our annual planning process, so that we can ensure wish work gets a dedicated spot in our roadmaps for the next fiscal year.
I took another look at this year’s timeline and still think it's best to keep the plan as-is, because we want to pick up tickets as early as February in this first cycle. Carrying out this entire new process in January/February instead of doing the submission and triage in Oct/Nov would push that back. Plus we don’t want to pull the vote to December, because we want everyone to get an equal chance to participate, but we know that a lot of people take a break then.
Also, there’s a similar discussion happening on the proposal Talk page in case you want to follow along there. SPerry-WMF (talk) 21:22, 1 September 2026 (UTC)Reply
@SPerry-WMF how does the wishlist interact with feature request tickets opened in phab? I usually just go the phab route. Do these just get lumped into one pool to be evaluated? Is one the preferred process over the other? RoySmith (talk) 13:34, 29 August 2026 (UTC)Reply
See phabricator as a permanent list of all that could potentially be done, but no promise of anyone even looking at it. Its also generally more technical. Wishlist is a subselection of that same list, but allows community voting, and comes with a promise of people actually evaluating what was filed. —TheDJ (talkcontribs) 17:36, 29 August 2026 (UTC)Reply
A well-functioning Wishlist will be able to get top wishes prioritized, with WMF teams assigned to work on them. On Phab, whether something gets prioritized is completely up to the team or volunteer devs that maintain the software. Also, Phab skews towards existing software rather than new software. As for which system a community member should use, maybe start by creating a Phab ticket, and then if the issue is important AND not getting worked on, also file a wish. Although this "strategy" part is subjective so is up to you. –Novem Linguae (talk) 10:21, 30 August 2026 (UTC)Reply
  • I have a proposal: I propose that multiple people who read this and agree with me place/support a wish on the wish list that reads something like this:
"Knock down our technical debt by fixing small quality-of-life issues, prioritizing things that are easy to fix. Don't hold back because fewer people are affected, because some unpaid volunteer supposedly maintains a script, or for any other reason. If it's wrong and you can fix it in a short amount of time, fix the bug wherever it resides."
Regarding "fewer people are affected" I have seen this as an excuse for not fixing an arbcom script that doesn't count words correctly (few people are the subject of an arbcom case) and as an excuse for not fixing a nasty accessibility bug (few Wikipedia editor are blind).
Probably best if someone else makes the wish. I have a number of people who hate me because I suggested that the WMF stop pointing a giant money hose at their pet projects. --Guy Macon (talk) 18:43, 4 September 2026 (UTC)Reply
@SPerry-WMF, when you were looking at wish "size", how small was "small"? Might there be room for "extra small" wishes? In solidarity, asilvering (talk) 20:31, 4 September 2026 (UTC)Reply
I would consider the smallest possible "fix a bug" job to be something that you give to a developer and they say that they can completely finish the job including documentation in three hours or less.
(I don't trust 15-minute fixes on anything someone else is supposed to use. Too many times the fix adds a new bug. I would say that devoting maybe an hour to testing is reasonable for the smallest, easiest bug. More if you do regression testing).
When you apply the standard multiplier (start by multiplying every estimate any developer gives you by Pi and then look for reasons it might take longer) you can hope to get it knocked out in about a day.
If they say they really can knock them out quicker that that (you might have managed to hire the next Charles H. Moore or Margaret Hamilton), have them do ten have someone else check them for errors, and you will have a better estimate than I could come up with.
The key is picking obvious quality of life issues that nobody is even thinking of fixing. Here is another example:
Without checking, tell me what happens if I sign this post with 1 tilde, 2 tildes, etc. up to 12. Can you? I can't without checking my notes. How long would it take to fix the stupidity that results from something that is 99% likely to be a typo? It will take some amount of time to decide how to fix it. Automatically turn everything from 3 to 12 into four tildes, screwing with the tiny percentage of users that want to add just a name or just a date? Tack on an "are you sure" and "never warn me about this again" dialog (my preferred fix)?
Just for practice in the proper method of fixing stupidities like this, refrain from telling me about all the cool ways that I can abandon the workflow I have been using for years and switch to a method that doesn't require the tildes. Dumping the bug back on the user and asking them to use a workaround might be the only answer you have, but it should never be your first choice. --00:13, 5 September 2026 (UTC)Guy Macon (talk)
The sizing we use currently is as follows:
  • S - not complex at all
  • M - mostly clear, but some complexity, solvable by one team
  • L - some complexity, potentially requiring help from multiple teams
  • XL - very complex, likely requiring multiple teams
There is definitely room to add an XS.
@Guy Macon: I think that's a fair suggestion, but what would make it really actionable is a list of phab tickets that you'd prioritize. I worry that with a wish formulated like that, it could be anything and everything and the impact is in the eye of the beholder. Plus we wouldn't ever be done with the wish, because there are so many tickets going back years, so one year wouldn't be enough to get through them all, even if we could put a full team behind it. And that's the other aspect that makes your proposal difficult to execute: wishes will be done by the teams who are best equipped to fulfill it. For a wish as you describe it, it would be pretty much every product team at the Foundation. What would get you more success is to go by general area or tool: Have one wish that lists tickets you want closed that address mobile web editing, one that fixes all issues you might see with the Watchlist, etc. That way, people voting on your wish would know exactly what all small changes/fixes you are referring to and you'd give them a tangible result for what would happen if they cast their vote on this and it would be worked on. SPerry-WMF (talk) 16:35, 11 September 2026 (UTC)Reply
Too much work, needs a full team to get done? So don't get done. Put one developer on it for three days, then next week put another developer on it for another three days. The important thing is steady progress instead of doing nothing to reduce technical debt.
Asking an unpaid volunteer to help you to decide what to work on? That's the kind of thinking that got you so deeply in technical debt. When something isn't working, don't do it harder. Let the developer decide what to work on. Make the only instruction to the developer "do a bunch of small things that you can knock out fast." Don't worry about importance, impact, or who "owns" the problem (fix it whether it is in the code you maintain or some script some volunteer wrote), or anything else. Just start someone plugging away at it instead of doing nothing. If it becomes obvious that the wrong things are being done we can address that later. If for some reason you decide that everything needs to be fixed in a year (you have been fine with zero effort to address small technical debt issues for decades) we can address that later.
Put up a simple web page where you list what got fixed, what you gave up on because it didn't turn out to be as easy as you thought, and things you are working on or thinking of working on. Let anyone make suggestions (most of which will be useless), but let the developer decide what to fix. They are in the best position to know what they can fix quickly. Keep it simple, Do something instead of doing nothing. --Guy Macon (talk) 17:35, 11 September 2026 (UTC)Reply
"Put one developer on it for three days, then next week put another developer on it for another three days." That's not how software is written or fixed. I doubt many of the tickets could be resolved by one person working for three days. But imagine, like, suggesting that a novel be written by having one author work on it for 3 days one week, another author continue working on it for 3 days the next week, etc. Patchwork code writing like this is a bad idea. (It's probably how a lot of the bugs and lack of future-proofing and such got in the code in the first place: too many hands stirring the soup.) Levivich (talk) 17:56, 11 September 2026 (UTC)Reply
I am intimately familiar with the way software is written and fixed at Boeing, Airbus, Mattel, Perkin Elmer, Parker-Hannifin, NASA, and a number of smaller companies that you have never heard of. What you say is perfectly valid for small to large software projects. It is absolutely not true for the kind of tiny, easy-to-fix quality of life issues I am talking about, and indeed most of the companies I just mentioned have someone quietly chipping away at the kind of small technical issues that "proper" software development has trouble dealing with. Like a word counting script, written by a volunteer who hasn't logged on in years, and which fails to accurately count words. That's about a three hour job including testing and documentation. --Guy Macon (talk) 18:24, 11 September 2026 (UTC)Reply
We're making decisions as part of the Foundation's annual plan. The wishlist is specifically meant to complement that, meaning the decision on what should get prioritized as part of wish work should sit with the community and will be voiced via vote count. Making clear what is and is not part of a wish is an important factor in that decision making process. So yes, in this case asking volunteers to decide on what to work on for a wish is exactly what we should do. SPerry-WMF (talk) 18:59, 11 September 2026 (UTC)Reply
I am going to withdraw from this discussion now, convinced that you completely failed to understand what I am asking, and assuming that the fault is on my end and that somehow I am incapable of communicating. Maybe someone else will have more sucess than I did.
What you are talking about is what you are doing now. I have no reason to think that what you are doing now isn't being done well. It's as if I asked someone "please take the garbage out. It will only take a minute." and they replied "I already have a plan to repair the leaky roof and getting three estimates is exactly what we should do." Not the same thing. You will no doubt end up accomplishing many new and useful things while making zero progress on the easily-fixed technical debt issues I am talking about. And everyone who participates in any Arbcom case will still have to deal with a word count tool that doesn't count words, sanctions if they go over, and clerks who eyeball the word counts and make estimates as a workaround for a shitty word-count tool that you could have fixed in a few hours.
So here is my final feedback: In my opinion you are so focused on doing what you are already doing quite well even better that you appear to be completely blind to the possibility of also doing something that is low effort unless you can approach it the same way Procrustes approached innkeeping. I am done. --Guy Macon (talk) 20:24, 11 September 2026 (UTC)Reply
SPerry-WMF is possibly saying that their hands are tied and directions from above mean that fixing things won't be done. WMF bureaucracy probably has learned Microsoft's mantra that time spent fixing things gives their competitors time to develop more glitz that will steal a lead. Levivich is talking about different things—items larger than what you have in mind. Guy Macon is absolutely correct that having one person focus on one small problem of their choice is exactly the way to make quality progress. If that person needs help because it's more tricky, try two people. After that, think about what to do. Make progress on technical debt. Johnuniq (talk) 03:20, 12 September 2026 (UTC)Reply

Wikimedia Foundation Bulletin 2026 Issue 16

[edit]

MediaWiki message delivery 20:59, 1 September 2026 (UTC)Reply

Unionization election results

[edit]

are available at this link: 158 votes for, 14 against. WMF statement here. Extraordinary Writ (talk) 23:57, 3 September 2026 (UTC)Reply

I wasn't able to view the first link (the page on the NLRB website) until I used a VPN to get a US IP address. The only other information that I found useful there was that there was a single void ballot (which is unfortunate for that individual if it was unintentional), and there were no challenged ballots which is at least a suggestion that all three parties (WMF, WWU, NLRB) regarded it as fair, although it hasn't been officially certified yet (according to the WMF statement, although the implication is they are expecting that to be just a formality).
There were 213 eligible voters and 173 ballots cast (including the void one) meaning the turnout was 81%. I don't know how that compares with other comparable elections (a quick google produced answers ranging from 40% to 90% for mail-in ballots in sources that were at first glance equally reliable) but it is certainly a democratic mandate, and despite some fears the majority of eligible staff were able to participate (and it is unlikely that lack of ability is the reason for every instance of non-participation).
I note the WMF explicitly respect the outcome and commit to engaging in collective bargaining in good faith. Do not expect instant results though. Even if both parties engage in the absolute best of faith, the differences are exclusively minimal and exclusively relate to areas that are simple to negotiate, it would likely take at least a couple of months after negotiations start (for logistical reasons that may not happen instantly) and at least some of the union's grievances relate to areas that are objectively not simple to change and so will take longer. Things not happening quickly is not evidence, in and of itself, of either party acting in bad faith. Thryduulf (talk) 01:15, 4 September 2026 (UTC)Reply
Given how many job titles there are alone it's going to take a while to negotiate and that's before considering that a bunch of the union demands are non-financial which will also require long negotiation. If it were done faster than a year I'd be surprised, even with both parties bargaining in good faith in ideal circumstances. I remain unsure that this union is good for editors, but I am sure that this result - nearly 3/4 of all eligible employees supporting the union - shows how if the foundation had learned from our !vote approach a lot of time, money, ill-will could have been saved/avoided. Best, Barkeep49 (talk) 01:49, 4 September 2026 (UTC)Reply
Just out of curiosity (and if you don't mind me asking), what makes you unsure that this union is good for editors? Some1 (talk) 02:31, 4 September 2026 (UTC)Reply
Because I look at what they originally had as their focus areas (which has only changed slightly today) and I see multiple places where what the community wants will come into conflict with what the staff wants. For instance around annual planning. The community has a hard enough time in the current process getting our voice heard. With a union staff will now have ways to ensure their priorities are heard and honored in a process which - because the foundation decided a Global Council could not happen - editors do not and will not have. Notably, they also acknowledge how the community and the union are not in harmony hence their asking for mental health support for WMF frontline workers (frontline workers are the people who interact directly with the Wikimedia movement communities). This got sanded down when the union realized the community was going to be really helpful in them achieving recognition and in order to get some staff who have a different view of the community on board and eager for the union. I genuinely hope that the union really does ❤️ the Wikimedia movement communities as they say in their approach, but even if they do they're no substitute for the community acting for ourselves. So yeah I remain uncertain that when the community isn't upholding our values and demanding the same of the Foundation Leadership and the Board in order to get the staff what they plainly deserve, that the union will stay in solidarity with us. Best, Barkeep49 (talk) 03:15, 4 September 2026 (UTC)Reply
I wouldn't put much stock in the mental health support comment; even if the community was consistently civil, humans aren't meant to have 20+ people express disapproval with something they've said. Fundamentally, that's just not how we work.
To use an example people are more disconnected to: let's take a school. Even when parents and teaches, and admin are all on the same side, dealing with parents/students is incredibly stressful; a psychiatrist I know probably ended up treating half the admin and teachers in the school district. That doesn't mean the teachers who need better mental health support somehow aren't in harmony with the students, or their goals conflict, it's just a fact of life. If I'm on my feet all day splitting wood, I need good shoes and gloves and regular water breaks to minimize serious injuries. If you're dealing with unhappy people all day, you need mental health support. GreenLipstickLesbian💌🧸 04:02, 4 September 2026 (UTC)Reply
I feel like every decent job should support the mental health of their workers, especially anything that deals with the public in any sense. Clovermoss🍀 (talk) 04:20, 4 September 2026 (UTC)Reply
Note that a large part of the desire for mental health support for front-line workers is for the trust & safety folks dealing with cases of harassment, threats of violence, threats of suicide, posting of wildly inappropriate stuff like griefing with gore and child sexual abuse material, etc. It's not about "an editor was rude to us" as such. brooke (talk) 17:56, 4 September 2026 (UTC)Reply
If a mental health support programme for staff dealing with harassment, threats of suicide or child sexual abuse material does get introduced as a result of this unionisation effort—and I sincerely hope it does (and am quite surprised it doesn't exist)—then as someone who has been dealing with this as a volunteer for well over a decade, I hope the Foundation could use this opportunity and offer a similar programme (or at least some related training) for functionaries such as oversighters or stewards. Thanks for mentioning this @brooke. odder (talk) 23:10, 5 September 2026 (UTC)Reply
I'd point out that this VP page has a giant civility warning at the top of it to in essence telling people not to verbally abuse WMF staff. The fact that such a warning is necessary shows that there have been problems in the past. I think supporting staff mental health is a good thing. Both for staff and the community, as mentally healthy staff can engage the community more productively. I also think the annual planning thing is probably good or at worst neutral for the community - i think staff closer to the ground have more ties to the community so their input is likely to be more aligned with the community. Bawolff (talk) 09:52, 4 September 2026 (UTC)Reply
As the person who originally pointed that out, I want to reiterate that my support for unionizing has nothing to do with whether "the community" is going to get anything out of it. Is that comment/plank/whatever hurtful? Yes, very much so. Does that mean unions in general and/or this union specifically are bad No. Gnomingstuff (talk) 13:38, 4 September 2026 (UTC)Reply
I don't think I'd say its hurtful. I absolutely love this site, but some patterns in how we, as a community, act has led me to taking breaks away from here. But WMF employees don't have that luxury. Furthermore, I think we sometimes fail to differentiate between WMF staff and management, and that definitely doesn't help things, mental health wise. MetalBreaksAndBends (One for all) 14:12, 4 September 2026 (UTC)Reply
As I wrote above, supporting the desires of the staff to unionize is in keeping with our collective values. That's true even if the negative case (rather than the positive case) materializes and it's bad for editors. That's why it's a value - we hold onto it even when doing so is hard. Best, Barkeep49 (talk) 14:32, 4 September 2026 (UTC)Reply
Unions always carry risks. Concerns would be the desire to keep failed projects going to keep people in jobs or spend down the reserves rather than reduce headcount when donations fall. Things like that. ©Geni (talk) 21:14, 5 September 2026 (UTC)Reply
IIRC the union reported ~70% of eligible workers voted in favor in the first two online ballots. 158/213=74% in the third unnecessary paper ballot. Levivich (talk) 04:54, 4 September 2026 (UTC)Reply
Yes. I hope they learn something from all that time and money when it comes to the UK and any other countries where staff decide they want a union. I could certainly see some place where there's a smaller level of support where a full process would be needed to determine true wishes. But that's not been the case with the two countries who've sought a union so far. Best, Barkeep49 (talk) 14:30, 4 September 2026 (UTC)Reply
Two quotes jumped out at me from the Wired article about unionization:

The foundation declined to voluntarily recognize the union, contending that a vote would be a more democratic process in line with its principles. But the foundation’s leaders stayed neutral through the voting process, workers say.

Workers reporting that "leaders stayed neutral" is a good sign. I take it to mean there were no further messages like m:WWUEMAIL in the run-up to the NLRB vote. To me, that's a sign of WMF leadership responding to the community's complaints.

Concerns about pay and benefits were not major motivators for the union drive, according to Wikimedia employees. Instead, they say, frustration over how the organization explained several hiring and firing decisions in the past year provided the final push the union needed to win the vote.

I take back what I said before about rich workers wanting to get richer. Apparently not! I'm no expert on labor unions, but it seems notable that workers are telling the press this isn't about money, it's about how management manages. Levivich (talk) 18:37, 7 September 2026 (UTC)Reply
On your last point, at least in the UK whenever a union goes on strike the public (and lazy journalists) assume it must be because they want more money. Sometimes that's (partly) true, but whenever other issues are the driving force (most commonly safety issues in the rail industry) the union have to work really hard to get that message across so it's good to see Wired listening to what their sources are telling them. Thryduulf (talk) 19:34, 7 September 2026 (UTC)Reply
Yeah I definitely associate union grievances with wages/benefits and safety conditions, my point of reference being blue-collar unions and public unions. But I've since done a bit of research, and apparently it's not uncommon for these newfangled tech unions to organize over management practices and a voice in operations, not wages/benefits (which I found less surprising once I thought about it). Levivich (talk) 22:17, 7 September 2026 (UTC)Reply
? Having jobs and keeping jobs sounds like the predicate wages and benefits issue. Alanscottwalker (talk) 12:45, 8 September 2026 (UTC)Reply

Has there ever been an example of unpaid volunteers unionizing?

[edit]

In most jurisdictions labor laws (including the National Labor Relations Act in the United States) explicitly deny union eligibility to anyone who does not receive or anticipate economic compensation as not being a statutory employee.

In 2015 Reddit found that while unpaid volunteers cannot form a union, they certainly can go on strike

When management of The Trevor Project attempted to silence unpaid volunteers and only allow paid union members to voice their concerns and frustrations, the union, Friends of Trevor United stood with the unpaid volunteers.

Many volunteer fire departments are operated by unpaid volunteers who cannot form traditional labor unions, Instead, many of them form Volunteer Firefighter Associations, who in many cases operate exactly like unions negotiating response protocols and equipment standards. Many firefighter's unions opposed this.

I think that our newly formed union -- AFTER reaching a labor agreement -- should form a "Friends of WWU" group that can present proposals and take not-binding advisory votes on them.

I don't think the WWU should support or oppose anything having do do with the relationship between the WMF and the volunteers prior to negotiating a contract. Best to make the official word "we don't comment on issues like this." --Guy Macon (talk) 06:23, 4 September 2026 (UTC)Reply

Love this question! Within the scope of US legislation National Labor Relations Act, this would not be possible. (Frequently NLRB challenges hinge on whether interns are...volunteer or working, in cases of academia, reality tv shows and other cases)
..but on other hand, loss of wages etc.. is not a principle concern of volunteer associations/unions. This is a topic that has come up frequently in different discussions and there is definitely appetite for it. I personally would like to explore this topic more after Wikipedia:WikiProject_Organized_Labour/2026_Online_Campaign is over. There are currently 788 editors registered from at least 12 language groups which would also bring us the cross-wiki legitimacy that enwiki is often criticized for (while being the most numeric). ~ In solidarity 🦝 Shushugah (talk) 10:29, 4 September 2026 (UTC)Reply
It was mentioned above that the foundation decided a Global Council could not happen. The WMF gets to decide no such thing. The WMF may (depending on the question above) decline to recognise an editors' organisation. However, it certainly can't stop one existing, making demands, and in the worst case backing up those demands with various forms of action. Certes (talk) 11:03, 4 September 2026 (UTC)Reply
This gives a lot of credence to the idea that fulfilling a labor activist role play fantasy has been a major factor in this whole saga. Thebiguglyalien (talk) 13:27, 5 September 2026 (UTC)Reply
One person's one comment gives a "lot" of credence to your theory about what's a "major" factor in a "saga" involving thousands... riiight... Levivich (talk) 13:37, 5 September 2026 (UTC)Reply
Thebiguglyalien, would you be so kind as to let me know what part of the post you replied to you are describing as "fulfilling a labor activist role play fantasy"? Thanks! --Guy Macon (talk) 15:37, 5 September 2026 (UTC)Reply
I’m trying to understand what the practical benefit of unionizing unpaid volunteers would be. What would such an organization actually enable volunteers to do that they cannot do through existing Wikimedia structures? And would those benefits be the same for volunteers in different countries, language communities, and Wikimedia projects, given that their interests and relationships with the WMF may be quite different? Golda Is In The House (talk) 19:19, 6 September 2026 (UTC)Reply
Let's pretend we had a "union" of Wikipedia editors who were willing to vote on a strike and stop editing if the vote passed. Now assume that a bunch of volunteers want the WMF to do something the WMF refuses to do, such as telling us when the strike busting Littler Mendelson law firm was first hired or telling us how much it cost the donors to pay for a randomly selected Wikimania. With a strong union we could force them to reveal those secrets.
If you want to see an example of something that affected Wikipedia editors in India but was of great interest to editors in the US, EU, Australia, etc. and could very well have triggered a strike, look at this:
Read about it here: Asian News International v. Wikimedia Foundation
Key quote: "the Wikimedia Foundation agreed to the court's request to disclose the identifying information of online users involved in editing the [Asian News International] page."
And there was nothing the volunteers could do about it. --Guy Macon (talk) 20:28, 6 September 2026 (UTC)Reply
And there was nothing the volunteers could do about it if there was a union or similar body either. Thryduulf (talk) 20:49, 6 September 2026 (UTC)Reply
I don't think that is true. Given the pretend world I created where you can get a significant number of editors to strike, The WMF would have to suspend editing, because the trolls and spammers sure wouldn't stop trolling and spamming. And that would have an affect on donations, interfering with whatever the hell they are spending all of that secret money on. --Guy Macon (talk) 21:07, 6 September 2026 (UTC)Reply
The key word in that comment is "pretend". In a pretend world you can do or not do anything you like. In the real world however we have to deal with facts. In your pretend world WMF's lawyers would be able to stick two fingers up to one of India's highest courts without consequence, in the real world things are not that simple.
I'm not interested in your labor activist role play fantasy as @Thebiguglyalien put it (although I wouldn't have used those words myself), I'm only interested in the real world. Thryduulf (talk) 21:41, 6 September 2026 (UTC)Reply
With all due respect, if you have already decided that unpaid volunteers unionizing cannot possibly happen in the real world, and you are unwilling to discuss hypotheticals that you believe cannot happen in the real world, I am at a loss as to why you commented on a thread about unpaid volunteers unionizing.
But for those following along who are willing to discuss a hypothetical, The WMF is perfectly free to tell the high court of India to shove their demands up their ass. Sideways. They have no jurisdiction in San Francisco. They already told the Supreme Leader of Iran to pound sand regarding Images of Muhammad. --Guy Macon (talk) 22:18, 6 September 2026 (UTC)Reply
I commented at first because I naively thought that you were talking about the real world having referenced a real-world court case (about which you seem to care naught about the detail). The WMF could tell the Indian court to get stuffed, but they could not do that without consequence. Please, actually take a moment to think things through and talk to people who understand how courts and things work before publicly accusing people of bad faith and whatever else. Thryduulf (talk) 22:42, 6 September 2026 (UTC)Reply
What are the eligibility criteria for joining this 'union', and how will the 'union' verify that members actually meet them? AndyTheGrump (talk) 21:39, 6 September 2026 (UTC)Reply
[1] Willing to join.[2] Willing to strike if the union votes to strike. --Guy Macon (talk) 22:18, 6 September 2026 (UTC)Reply
So no verification at all. Basically, anyone can vote, as many times as they like. The trolls will love this... AndyTheGrump (talk) 22:30, 6 September 2026 (UTC)Reply
Oh, identification! I thought you were talking about qualifications. SecurePoll and Scrutineers. --Guy Macon (talk) 22:44, 6 September 2026 (UTC)Reply
Anyone can of course form any group they want and set their own guidelines for unified action within the group. Without an employee/employer relationship, it wouldn't be a labour union, but the group could still co-ordinate to take unified action nonetheless. isaacl (talk) 22:34, 6 September 2026 (UTC)Reply
They can. But pretending to be a 'union', and pretending they are 'striking', achieves nothing at all. If there is a specific issue with something the WMF does, people can try to organise an editing boycott, but participation in that is going to be an individual choice. Not something that needs union-roleplay. Calling yourselves a 'union' does nothing but make yourselves look silly. It will have no legal basis, and the WMF will be under no obligation to take the slightest notice. They aren't going to negotiate with role-players claiming to represent an unverifiable number of contributors. A boycott might possibly have an effect, if participation is high enough, but that is going to be over single issues. AndyTheGrump (talk) 23:17, 6 September 2026 (UTC)Reply
Sure; that was implied by my saying it wouldn't be a labour union. isaacl (talk) 00:01, 7 September 2026 (UTC)Reply
We're already in a union. It's called "the volunteer community." If we want to strike, we strike. If we want to black out the main page, we do that. Etc.
Any subset of this group would be powerless. If, for example, 50% of active editors joined a "union," made demands, and went on strike, it wouldn't mean anything because the other 50% would keep working. Things might get slower, but half the editors won't be able to "shut it down" or get consensus to blackout the main page or post banners or anything like that.
And realistically, getting 50% to join anything would be near impossible. 1,000 editors in a union won't make much of a difference; the others will run the site during any strike. In order for boycotts or strike actions to work, you need much closer to 100%. Subgroups won't work. That's why it's only the entire collective that can be the union, and so it already is.
Which is good news. If you agree with me, then the union you want already exists! Huzzah! Now what? :-) Levivich (talk) 23:56, 6 September 2026 (UTC)Reply
I fully agree that getting enough editors or even admins to agree to any sort of collective action is practically impossible. It's an interesting thought experiment, but I doubt that it would get significant support even if WMF top management decided to do what Firefox decided to do long ago (Firefox has so far accepted a half-billion dollars from Google, and roughly 85% of Mozilla's overall revenue comes directly from Google).
However, for those who say that it is actually impossible I would point out that Reddit moderators did it (and got kicked in the teeth with classic 19th century railroad-Barron union busting techniques -- no legal protection for you!) and that it is an easily verifiable fact that there exist multiple Volunteer Firefighter Associations who operate exactly like unions, negotiating working conditions and equipment standards. So it isn't actually impossible. Just so unlikely as to be virtually impossible.
Because someone is sure to ask, by comparison to Firefox, The WMF gets about 4% of its revenue from Wikimedia Enterprise, Google was the first paying customer for Wikimedia Enterprise but now shares that 4% with Amazon, Meta, Microsoft, Mistral AI, and Perplexity, so Google is a fraction of that 4%. See meta:Overview of Wikimedia Foundation and Google Partnership. Feel free to ask the WMF exactly how much they get each year, total, from Google. They might even respond if it is as tiny a drop in the total annual funding bucket that I think it is. --Guy Macon (talk) 02:53, 7 September 2026 (UTC)Reply
You're trying to draw a straight line through the fight against 19th century railroad monopolies, working conditions for volunteer firefighters, people who stopped posting on Reddit because of some petty internal drama, and whatever self-destructive grandstanding we're now doing here. I don't even know where to start with that. Thebiguglyalien (talk) 04:17, 7 September 2026 (UTC)Reply
I am going to ask you nicely to please stop engaging in personal attacks. You behavior in inappropriate. WP:NPA. --Guy Macon (talk) 04:46, 7 September 2026 (UTC)Reply
A comment is not a personal attack because you don't like it. And you don't exactly have much ground to stand on with "inappropriate" behavior when you're making credible threats to harm the entire project. Thebiguglyalien (talk) 04:50, 7 September 2026 (UTC)Reply

A new village pump page?

[edit]

I think it would be a good idea to have a Wikipedia:Village Pump (affiliates). Sometimes affiliate related issues are shared here because there's some overlap, but I think it could be its own topic. I think it could also help people provide feedback on projects and it might encourage people from affiliates to share updates. I'll create it if other people think this is a good idea as well. Clovermoss🍀 (talk) 10:25, 8 September 2026 (UTC)Reply

I'm wary of adding yet more Village pump pages, and I've not seen much of a problem from them being posted here or at other relevant village pumps. Have there been problems I've missed noticing? Or call from people for a dedicated page (versus "might encourage")? Anomie 11:41, 8 September 2026 (UTC)Reply
Technically I'm a contractor for Wikimedia Canada and I think it's a good idea. But it's more of a might encourage thing, since I don't like that most of this communication takes place off-wiki in places like mailing lists/telegram. Encouraging more on-wiki communication promotes understanding from both sides and hopefully leads to people feeling better about some of their concerns since they have a place to go. Clovermoss🍀 (talk) 11:45, 8 September 2026 (UTC)Reply
I also don't see the need for another VP page. If there was so much traffic about affiliates that it was clogging up the existing pages, it would make sense. If the people involved with affiliates prefer to use other channels, it's not our place to say, "No, we want you to have those discussions here". Also, this isn't an enwiki-specific thing. I would think someplace on meta (i.e. meta:Wikimedia movement affiliates) would be a better place. RoySmith (talk) 12:04, 8 September 2026 (UTC)Reply
It's not a have thing, but an option. The people I talk to aren't trying to hide things and would like the community to like them. I think Meta isn't a bad idea, but a lot of things that affiliates do affect en-wiki specifically and it makes sense to have a page here for that as well. People applying for grants are also encouraged to seek feedback and those pages don't seem to have as much traffic as a specific page for affiliates and grants could. Clovermoss🍀 (talk) 12:11, 8 September 2026 (UTC)Reply
My first reaction is to explicitly add affiliates to the scope of this page. If turns out that this increases the traffic so much that WMF-exclusive stuff gets drowned out then we can split at that point. Thryduulf (talk) 12:48, 8 September 2026 (UTC)Reply
That makes sense to me. RoySmith (talk) 12:54, 8 September 2026 (UTC)Reply
It was my thought too. Per Thryduulf on rethinking if WMF stuff gets drowned out. (Slightly reduced risk of a dedicated throw brickbat space.) CMD (talk) 12:57, 8 September 2026 (UTC)Reply
I worry that might cause some friction, the way we tend to get upset when volunteer editors are conflated with the WMF. They're separate organizations, even if they're mission-aligned. I also envision encouraging a bunch of affiliate-related people to introduce themselves if such a page was started. Clovermoss🍀 (talk) 12:57, 8 September 2026 (UTC)Reply

I'm trying to understand the objections here. Is the concern that the page would end up so low traffic as to be dead? It's hard to know that without even trying. Or is it more that the uncertainty isn't great when there's already a bunch of village pump pages? Clovermoss🍀 (talk) 14:36, 8 September 2026 (UTC)Reply

onwiki issues related to affiliate related work, at least for my region, ESEAP, had been an once a year occurrence. Extending further out to the wider Asian region, anecdotally, had been one or two issues a year. I would prefer to utilise the misc village pump first before opening another venue. – robertsky (talk) 15:09, 8 September 2026 (UTC)Reply
@Robertsky: your affiliate region wouldn't have any ongoing projects or updates to share more than once a year? I'm not trying to dismiss your perspective, I'm just having a difficult time trying to understand it. There's no contests or challenges or institutional partnerships? Clovermoss🍀 (talk) 15:13, 8 September 2026 (UTC)Reply
You're a contractor and you want to create a page to benefit your contract? Alanscottwalker (talk) 15:20, 8 September 2026 (UTC)Reply
No, it doesn't really have anything to do with that. I'll be doing what I've been doing regardless of whether this is created. Clovermoss🍀 (talk) 15:22, 8 September 2026 (UTC)Reply
I thought you said you were a contactor of an affiliation and you thought this new page would be a good idea? Alanscottwalker (talk) 15:25, 8 September 2026 (UTC)Reply
I'm a contractor, but my contract has to do with running events and trying to increase the reach of Wikimedia Canada. Almost all of that stuff is offline, although I have started working on a subpage for transparency's sake at User:Clovermoss/Wikimedia Canada. A lot of stuff that goes on at Wikimedia Canada happens without me, as they have dedicated full-time staff. I made the distinction because I am not that. I have had an interest in making affiliates less of a maze to people outside of them for awhile, though. See this essay. Clovermoss🍀 (talk) 15:29, 8 September 2026 (UTC)Reply
The purpose of this new page would be to "increase the reach" of affiliates and "make affiliates less of a maze"? Alanscottwalker (talk) 15:37, 8 September 2026 (UTC)Reply
I doubt it'd do much to increase the reach per se, as again, most people would be whatever they're doing without that page. My desire to create something like this comes from the years I've spent not being involved with affiliates as a "regular" volunteer and wanting to understand how they work. Most of this information is very difficult to find if you are not part of those circles. Clovermoss🍀 (talk) 15:39, 8 September 2026 (UTC)Reply
How is it not increasing their reach when the page's purpose is to expose others to them? Alanscottwalker (talk) 15:50, 8 September 2026 (UTC)Reply
Because when I'm expanding the reach, I'm doing that in-person with Canadians who have incredibly limited knowledge about Wikipedia and Wikimedia projects. What I'm proposing here has more to do with people having a place to go that doesn't involve travelling to a conference to talk to affiliate-related people and for affiliate-related people to ask for feedback from people who may have very different experiences then their own. Wikimedia Canada is only one organization of several, let alone all these groups. Your concern seems to be based along the lines of someone told me to do this and no one told me to do this. Clovermoss🍀 (talk) 15:59, 8 September 2026 (UTC)Reply
No, I'm not concerned about what you are told. You have a contractual interest in an affiliate, and the purpose of this proposed page is for affiliates. Alanscottwalker (talk) 16:05, 8 September 2026 (UTC)Reply
I guess, but that's a cynical way of looking at it. The most personal gain I might get out of this is brownie points, but it's more of a risk than a benefit because it probably wouldn't look that great on me if it made people more skeptical of affiliates instead of restoring some trust.
I'd be interested in this even if I had nothing to do with Wikimedia Canada (as was the case a few months ago) but obviously I can't force you to believe in my sincerity. I mainly brought it up because I was being perceived as an "outside" source and people seemed to have concerns about if I was making assumptions about what spaces people connected with affiliates might be interested in. Hence my technically response to Roy. Clovermoss🍀 (talk) 16:15, 8 September 2026 (UTC)Reply
It is not cynical, it is standard concern. (see, WP:EXTERNALREL) Alanscottwalker (talk) 16:20, 8 September 2026 (UTC)Reply
I know what external relationships are. It doesn't undermine the primary goal of improving the encyclopedia. These two things aren't in conflict with each other. So yes, I think your take is fairly cynical even if I won't tell you you can't have that opinion. The fact that people see affiliates as being an external relationship speaks volumes in its own right. I think a lot of people who are more connected with affiliates would find that sad and want to change that perception. But they can't do that if they don't know who they even have to convince. Clovermoss🍀 (talk) 16:23, 8 September 2026 (UTC)Reply
Affliates are designed to be and have always been external because affiliation is not English Wikipedia's WP:PURPOSE. English Wikipedian's do not pay anyone to do anything, for anything, they are required to join no groups. I think you are undermining the encyclopedia, if you can't keep your external relationships, like your contracts, where they belong, not here. Alanscottwalker (talk) 16:43, 8 September 2026 (UTC)Reply
If you seriously think that, please start a noticeboard thread instead of casting unfounded accusations about me undermining the encyclopedia. Clovermoss🍀 (talk) 17:56, 8 September 2026 (UTC)Reply
It's founded in your contract, and in your refusal or inability to recognize that as an external relationship. Alanscottwalker (talk) 18:03, 8 September 2026 (UTC)Reply
Would you say the same about someone who worked for the Wikimedia Foundation? Clovermoss🍀 (talk) 18:06, 8 September 2026 (UTC)Reply
Apart from the limited office action, they are not to create things on this site, and they don't. Moreover, unlike any other corporation or group, they are the webhost and define the terms of use. Alanscottwalker (talk) 18:11, 8 September 2026 (UTC)Reply
The Wikimedia Foundation creates things that affect the English Wikipedia, alongside other projects, all the time. As well as commenting enough for this noticeboard to exist. I don't understand why you think it's disruptive to want a place to talk to people, when it's completely harmless. To say that I'd be undermining the project for doing so is confusing if you're not just trolling me. But I'd like to think maybe you just don't understand what affiliates are intended to do. Clovermoss🍀 (talk) 18:17, 8 September 2026 (UTC)Reply
Harmless? It is well recognized that editing with a conflict maybe harmless (it is also well recognized that it is not commentary on anyone's good faith or competence), it is equally well recognized that it needs to be avoided, and in talks, disclosed. Alanscottwalker (talk) 18:26, 8 September 2026 (UTC)Reply
AGF Czarking0 (talk) 16:27, 8 September 2026 (UTC)Reply
I appreciate the sentiment, but I can take a bit of pushback. I'd rather people voice their concerns than keep them to themselves. I don't want to dismiss people based on tone and cynicism alone because that drives me up a wall when I'm on the other end of it. Obviously I can't force other people to see my heart, but I hope it comes across. Clovermoss🍀 (talk) 16:31, 8 September 2026 (UTC)Reply
Well, since that didn't end up working out and this thread appears like it came after the above if one isn't looking at the timestamps, I'll note that I did end up escalating this to a noticeboard thread. Clovermoss🍀 (talk) 19:06, 8 September 2026 (UTC)Reply
While being a contractor might influence your perception of the ask, frankly it should not be much of a factor. I can see a potential for such a page to be a centralised location for communications between editors and the various affiliates who have activities that touches on this project. Not all editors are comfortable reaching out to affiliates offline, and affiliates do not have a monopoly of ideas of what activities to run for the project, and the editor in question might just need some support somewhere somehow to get their proposed activity off the ground.
That being said, while I am in the Singapore user group as well as the ESEAP Hub (as a steering committee member) and I see some merits in having one such village pump, I would like to reiterate that the we can potentially use the misc village pump first to see how the conversations play out, in terms of the frequency, breath and depth. If it gets too intense there, an affiliate village pump might then be created. – robertsky (talk) 16:05, 8 September 2026 (UTC)Reply
Does not WP:Affiliates suggest they are all online?Alanscottwalker (talk) 16:16, 8 September 2026 (UTC)Reply
@Robertsky: What about something like "an affiliate corner" in userspace for a few months (User:Clovermoss/Affiliate corner) just to see if there's enough to justify a "real" page? I think your idea has merit. If mine doesn't, everyone can just go back to Village Pump (Misc). If my idea actually ends up encouraging conversation, its text and history can be moved to a Village pump subpage for affiliates specifically. The main reason I want it all together is so people interested in those issues don't have to wade from unrelated conversations in archives to find them. Clovermoss🍀 (talk) 16:21, 8 September 2026 (UTC)Reply
If you started it in Wikipedia space, I'm doubtful anyone would bring it to MfD. It just wouldn't be a Village Pump during that time. Best, Barkeep49 (talk) 16:42, 8 September 2026 (UTC)Reply
Well, Wikipedia:Affiliate's corner it is then. Better than userspace, at least. I'll set it up later today when I'm less busy. Clovermoss🍀 (talk) 17:53, 8 September 2026 (UTC)Reply
There's already Wikipedia:Meetup. It probably makes sense to have both of those link to the other. RoySmith (talk) 17:58, 8 September 2026 (UTC)Reply
Yes, I agree. Not all meetups are organized/supported by affiliates, but some are, and it makes sense to link them together once the page is set up. Clovermoss🍀 (talk) 18:00, 8 September 2026 (UTC)Reply
we have had wiki loves Ramadan this year that bubbled up onto here or ANI(?). Last year there was a translation project that went awry (contained at AfC). There may be another translation project soon, different group running. There are institutional partnerships, for enwiki specifically though, mostly would be the Australian, New Zealand, Singapore, and possibly Philippines that may have them. – robertsky (talk) 15:21, 8 September 2026 (UTC)Reply
I think Hawkeye7 does some stuff with Wikimedia Australia. Tamsin and Giantflightlessbirds are who I think of when I think of Wikimedia New Zealand. There's also some Wikimedia NYC people I know of like Pacita (WikiNYC) and Pharos. Clovermoss🍀 (talk) 15:27, 8 September 2026 (UTC)Reply
Tamsin and Giantflightlessbirds did m:GLAM/Wikifying a Conference recently, and produced a guidebook/book on the topic, which should actually help affiliates with offline outreach. 😏 – robertsky (talk) 15:47, 8 September 2026 (UTC)Reply
I got a physical copy when I attended wikimania. :) Looks interesting, although I haven't got a chance to take more of a detailed look yet. But I wouldn't have really known any of the names I just pinged except for Hawkeye if I hadn't started attending conferences in the last 3 years and I'll never forget how completely obscure everything felt before then. Even now, there's still a lot of moving pieces I'm trying to make sense of. Clovermoss🍀 (talk) 16:02, 8 September 2026 (UTC)Reply
(Just noting that Tamsin from NZ is actually DrThneed.) Giantflightlessbirds (talk) 20:10, 8 September 2026 (UTC)Reply
Oops, thanks for the correction! Clovermoss🍀 (talk) 20:13, 8 September 2026 (UTC)Reply

I have three at least partially competing reactions to this: a logistical reservation, a pessimistic view, and an optimistic view. Logistics: yet another discussion page? I'd like to see it grow in an existing discussion page first. And most affiliates aren't even primarily active on the enwp, so agree with the above that it seems more like a meta page? For those that affect enwp directly, perhaps creating a signpost series that solicits regular updates from affiliates is a good place to start? Pessimistically, it's hard not to draw a comparison to this board, which is functionally less about WMF sharing things and more about a place to evaluate and air grievances about the WMF and/or what it's doing. Putting aside the extent to which that's needed and/or useful, "do for affiliates what this page does for the WMF" doesn't seem appealing for affiliates to participate in. We should remember that while there are a few well resourced affiliates the overwhelming majority of affiliate members are volunteers. The optimistic view: it seems like a lot of people don't really have a good idea of what affiliates are, who participate, what they do, etc. On pages like this one, they're more like a line item in a budget. On Meta they're often lists of metrics. Having a space for "here's some neat stuff we've been doing at [affiliate] that you might be interested in" could actually connect the community with the good being done. Like how many people here know that Habst and some other folks at WikiNYC have been developing a neat tool called Wikinewsie that follows edit-based news (articles that are being edited a lot, or by a lot of people)? So yeah, mixed feelings, tending towards a "this is maybe a signpost thing". Rhododendrites talk \\ 16:54, 8 September 2026 (UTC)Reply

Perhaps, if someone wishes to write a signpost article or column ('what X group is doing'), but it is not hard to find what https://meta.wikimedia.org/wiki/Wikimedia_New_York_City is up to or who to contact. Some people may be confused that they are a separate corporation with their own mission and that volunteering for them is volunteering for that corporation, but it is not hard to find that out. Alanscottwalker (talk) 17:22, 8 September 2026 (UTC)Reply
  • I admit skepticism, but I could be convinced. What exactly would we be talking about on an affiliate noticeboard? Do we have any examples of recent noticeboard conversations about affiliate stuff? As far as I know, affiliates aren't exactly in the business of writing or developing for EnWiki. They organize events, train people, try serve as centers of political power, channel money into... whatever it is they use all that money for, and run various competitions. But content? Tools? The sorts of things that affect our lives on EnWiki? I mean, for the most part, they leave us alone, and we leave them alone, and we're all the happier for it. Am I missing something about the work that affiliates do or that we should be talking about here? CaptainEek Edits Ho Cap'n! 22:39, 8 September 2026 (UTC)Reply
    Well, certain affiliates like Wikimedia Deutschland take on technical projects that the WMF doesn't. WPMED does as well (Doc James would be able to talk more about some of what they do). WikiPortraits takes photos that are used quite frequently on the English Wikipedia (SuperHamster could give more accurate stats than I could off the top of my head; noting that I also have done work with WikiPortraits for the purpose of this conversation). There's a lot that could be discussed and talked about. I also think it'd be good just in general for both sides to understand each other better. Like, one of the most common volunteer concerns is when things like contests go wrong and result in massive clean-up efforts. I also see interesting things that happen on wikimedia-l that could give people a better impression of the "good" side. For example, this recent thread about cyclists. It's easier to build trust if people know what's going on. Clovermoss🍀 (talk) 23:48, 8 September 2026 (UTC)Reply
    There's also a Commons user group , where doing stuff on-wiki is kind've the entire point. Some projects have much stronger overlap between affiliates/the general editing community. It's mainly North America that's the exception to that, where affiliate activity is more concentrated to specific areas. Wikimedia NYC does a lot of good work, though. Clovermoss🍀 (talk) 00:10, 9 September 2026 (UTC)Reply
    @Clovermoss Admittedly, my regard for WPMED is low; you can see my thoughts about its value at . I don't disagree that Affiliates can do good for the movement, but I'm less certain they're doing EnWiki relevant work? Or that is to say, work that wouldn't be more relevant at say Meta. I do appreciate you giving some examples though. CaptainEek Edits Ho Cap'n! 04:14, 9 September 2026 (UTC)Reply
    I agree it would be helpful to have some specific examples drawn from the past of what would be useful to have on a separate project space page. isaacl (talk) 01:28, 9 September 2026 (UTC)Reply
    Another example is trying to understand what concerns affiliates have about the new proposed funding model from the WMF. Sohom Datta knows more about that from me. The stuff that's being proposed sounds reasonable to me at a glance, but I don't know what people's objections are, just that they have them. A place where people can share in their concerns, hear other people's perspectives etc can be really valuable.
    Another reason is that affiliates do have some forms of power in the movement. For example, affiliates were given the ability to shortlist which candidates were able to make it to board elections last year. Bluerasberry can probably give more examples than I can as they've been immersed in that world for longer. Affiliates can also go hand-in-hand with WikiProjects. For example, Wikimedia LGBT and WikiProject LGBT. Clovermoss🍀 (talk) 01:33, 9 September 2026 (UTC)Reply
    Are you suggesting that affiliates are interested in posting on English Wikipedia about their concerns with the proposed WMF funding model, and about what candidates they are approving to proceed to the board elections? (I might not have been clear regarding what I meant by examples drawn from the past: are there past messages that were posted elsewhere on English Wikipedia that could, in future, be posted in the new proposed venue?) isaacl (talk) 02:20, 9 September 2026 (UTC)Reply
    Yes to the first, uncertain as to the second unless we're talking about cleanup threads that happen at ANI and the Wikipedia:Education noticeboard. A place for good things might help people feel like there's less of a disconnect in priorities.
    Most of the communication I'm hoping for to take place exists in places that are not on-wiki, like an affiliate's website, offwiki Telegram groups, in-person, Whatsapp, etc. I'm just hoping to convince some people who are amenable to the idea to participate on on-wiki discussions if they wish to, but I want to emphasize the importance of optional and respecting people's autonomy to not participate if they do not want to. Clovermoss🍀 (talk) 02:40, 9 September 2026 (UTC)Reply
    They're your examples, so you can tell me what kinds of messages you think they would cover. Your second example was "affiliates were given the ability to shortlist which candidates were able to make it to board elections last year," so I'm not sure what you mean by cleanup threads.
    I'm not sure why an affiliate holding conversations on their own web sites or various messaging apps would decide that a better replacement would be a page on English Wikipedia. I do appreciate that an affiliate who wants to hear more from the broader communities would want to hold discussions on wiki. But as Rhododendrites said, affiliates are largely composed of community members. I think most of them should be able to find places to start conversations on wiki. isaacl (talk) 02:56, 9 September 2026 (UTC)Reply
    I'm not trying to replace anything, but offer a supplemental place to people who want it. I disagree with Rhododenrdrites view, but I don't want to single out people who don't have connections with any community, but plenty of people like that do exist.
    This isn't a longstanding issue of contention for no reason, even if it's clear people have had very different experiences with different affiliates. I'm glad Rhododendrites has only ever seen that overlap, as there tends to be less issues when everyone is on the same page, hence why I wish for there to be a place for people to communicate more often.
    I've been giving lots of examples of what a page like this could cover. I don't think these issues contradict each other, they're just examples of different things affiliates could talk about with people who exclusively stay on-wiki. Clovermoss🍀 (talk) 03:11, 9 September 2026 (UTC)Reply
  • Having thought through this further, I am more cautious due to the way interest groups generally interact on English Wikipedia, which is through WikiProjects. Participation, and even offline participation, in Wikipedia:WikiProject Women in Red (WiR) is larger than participation in many (most?) affiliate projects. WiR has created an ecosystem of on-wiki pages that document their work, but doesn't have a dedicated board to discuss its activities with WP:MILHIST. Questions following from that include: would setting up like a WikiProject work similarly for affiliates? Are WikiProjects are also missing out on some communication opportunity? Do we want to treat mostly offline groups (eg. affiliates) differently from mostly online groups (eg. WikiProjects), and if so, why (and how to deal with the sliding spectrum between those)? CMD (talk) 23:18, 8 September 2026 (UTC)Reply

A lot of "build it and they will come" venues get proposed, which is why often there's a reluctance to create yet another page that no one reads or edits. As Barkeep49 said, though, I don't think anyone will make a fuss about a new project space page, as long as it doesn't impose any additional work or mandatory changes to workflow for anyone not interested in using it. This does mean that affiliates shouldn't expect that any message they place in this new venue will get read by a broad segment of the commumity. So if they are seeking broader awareness, they still need to use one of the existing appropriate venues to garner the desired attention. isaacl (talk) 01:24, 9 September 2026 (UTC)Reply

I think part of the issue is that it isn't always clear what the appropriate venue is, which is why I suggested making miscellaneous affiliate communication stuff part of this boards focus (could also be a different existing board of course). If input is sought regarding a project with a specific topical or geographical scope there may be existing noticeboards for that, but there isn't anywhere (obvious to me at least) for projects without that focus and/or where there isn't an existing board closely aligned. Thryduulf (talk) 01:47, 9 September 2026 (UTC)Reply

I will also note, though, that I'm wary of creating a page under the assumption that others want to use it, if we haven't heard from the others. Is there a clamor from any affiliates that a separate venue for them is needed? I think we should consult with them first and ask if they want a new venue. Plus it feels very English Wikipedia-centric to assume that affiliates want to communicate with the English Wikipedia community. isaacl (talk) 01:33, 9 September 2026 (UTC)Reply

I could ask other people, but I'd worry I'd get accused of canvassing. I was mostly just going off the impression I've had in off-wiki conversations that affiliate-related people aren't inherently opposed to engaging with the community more. A lot of people see themselves as part of the community and don't like the idea of being perceived as some seperate, outside force. I think the main concern would be people treating them like Alan just treated me here. It's a bit hard to convince people to go participate in an optional place where people might accuse you of doing all sorts of things you're not doing. But I'm from the on-wiki side of things first, so I understand the importance of constructive criticism and giving people an outlet. It's difficult to rebuild trust by doing nothing. Even just identifying that a problem exists can be a helpful step. I've found certain people have been surprised by Wikipedia:Editor reflections, particularly my on-going analysis when it comes to what people think about WikiEd. Clovermoss🍀 (talk) 01:49, 9 September 2026 (UTC)Reply
Also, having a space on enwiki doesn't mean there couldn't be pages on other projects where affiliates engage with the relevant communities. I think we're the only project with a dedicated WMF page (and am not sure what led to the creation of it), but other projects could theoretically build their own place for this if they think it is a useful concept. Clovermoss🍀 (talk) 01:52, 9 September 2026 (UTC)Reply
See Wikipedia:Village pump (proposals)/Archive 168 § Proposal: New Village Pump Page. This page was created with the goal of increasing communication with the WMF. Their communication staff, though, didn't commit to post everything to this village pump page, as it didn't want to split discussion if another page was more appropriate. isaacl (talk) 02:11, 9 September 2026 (UTC)Reply
Thanks for the link! It looks like Alsee did good work on getting that proposal through. I wonder if they agree or disagree about whether this is a similar sort of situation nessecitating the need for a new page. I feel like their experience is so 1-to-1 here in a way you rarely get when it comes to proposing new things. Clovermoss🍀 (talk) 02:19, 9 September 2026 (UTC)Reply
I appreciate there are some who like the "build it and see if they come" approach. Personally, I prefer to gauge demand on both sides (those who would potentially edit the page and those who would read it) and figure out if it is sufficient to sustain a new venue. I feel that for a new venue to work, it needs to be publicized and some effort invested into making it a minimally useful forum, but if there's no interest in using it, I don't want to push people into it.
I imagine the vast majority of affiliates don't have a lot of spare personnel for outreach to many Wikimedia communities. So asking them to come to English Wikipedia in addition to meta, the more typical cross-Wikimedia site, feels to me a bit like asking for special treatment. isaacl (talk) 02:02, 9 September 2026 (UTC)Reply
If they'd prefer to engage on Meta with their limited time, that's their choice, but that's not my preference. I'd be willing to take an active role at Wikipedia:Affiliate's corner but not on Meta due to some rather upsetting recent experiences. I have more faith in our processes for dealing with problematic behaviour, even if they aren't perfect. Clovermoss🍀 (talk) 02:15, 9 September 2026 (UTC)Reply
I understand your position. But I don't think that means we should tell all affiliates that the English Wikipedia community has more faith in its own processes, so they should come here for any discussions they want to have about their initiatives. isaacl (talk) 02:25, 9 September 2026 (UTC)Reply
But that's not what I'd be telling people. If people have more faith in Meta, they should participate where they feel the most comfortable. I just personally wouldn't be interested in having these conversations there. Clovermoss🍀 (talk) 02:31, 9 September 2026 (UTC)Reply
Well that just goes back to my question: why are we asking affiliates to engage specifically with the English Wikipedia community? Wikimedia Deutschland of course has a specific interest when it is working on a MediaWiki feature that is of interest to English Wikipedia, and it has indeed come to existing venues to discuss its work on improving the autogenerated citation names, without needing a new venue. But for many others, it feels like we'd be asking them to do something special for English Wikipedia. isaacl (talk) 02:46, 9 September 2026 (UTC)Reply
My understanding is that WCNA is often seen as somewhat of an English Wikipedia gathering, which may provide a different perspective to affiliates elsewhere. CMD (talk) 03:17, 9 September 2026 (UTC)Reply
WCNA is technically a user group in its own right, even if it does have a strong enwiki focus. Clovermoss🍀 (talk) 03:39, 9 September 2026 (UTC)Reply
Well that's confusing! I meant it in the, uh, sensu lato sense of groups regularly attending the conference. CMD (talk) 06:32, 9 September 2026 (UTC)Reply
What is this affiliate-related people aren't inherently opposed to engaging with the community more stuff? (this is also partly a response to a couple other comments). Affiliate-related people are part of the community. There are Wikipedians who don't engage with anything affiliates do, but if there are affiliate members that don't engage with one or more of the Wikipedia projects and their communities I haven't met them. Affiliate people are both existing volunteers who say "hey, maybe I'll go to one of these in-person meetings" and readers who become volunteers through the in-person meetings. I get there's a tendency to frame affiliates as some sinister cabal, but they're just you (the general you) if you went to an event, or took part in an edit-a-thon run by an affiliate, etc. And if we're specifically talking about people who get paid through an affiliate, we're talking about such a teeny tiny portion of "affiliate people" that there's not much to talk about. Rhododendrites talk \\ 02:20, 9 September 2026 (UTC)Reply
It is now over ten years since my time with an affiliate came to an end, so I think I have sufficient detachment to comment here. Yes affiliates are part of the movement, many people involved in affiliates were wikimedians before during and after their time with an affiliate. I like to think that applies to my time at Wikimedia UK, and many of the people in GLAM roles that I met in other chapters. But there are people who work for affiliates who need a little guidance and maybe some targets to interact with the volunteer community. I'm not convinced that the village pump of the English language Wikipedia is always the best place for that, The GLAM newsletter, The Signpost, Meta, outreach wiki if that still exists, are all relevant places/media depending on the topic. I do think that the WMF KPIs for affiliates should include some targets for interaction with the most relevant volunteer communities, and village pumps would be part of that. However a separate noticeboard should only come after the amount of postings has reached a point where some on the main noticeboard would like those threads spun off to somewhere they can choose not to follow, not before. Pave the desire lines, don't try to predict where the desire lines will go. ϢereSpielChequers 13:39, 9 September 2026 (UTC)Reply
@WereSpielChequers: That sounds more like a directory than a place for conversations, which isn't a bad idea in its own right. I'm just not sure how one would pave the desire lines without fragmenting discussions? Could you elaborate a bit more on what you envision here? Clovermoss🍀 (talk) 16:42, 9 September 2026 (UTC)Reply
Hi Hannah, well I think I am suggesting fragmenting discussions. So for example, when an affiliate is running a GLAM program that is relevant to a particular Wikiproject, I would hope that someone from the affiliate would talk on the Wikiproject page. If someone wants to discuss chapters in general then I think that discussion is best on Meta. Where I would like to see a change is with the individual Wikis that each chapter seems to have. At least they did when I was last involved in a chapter circa 2015. Having separate Wikis under the control of individual chapters might help with any proof if twere needed that those chapters were independent of the WMF. But it costs money and or requires extra volunteer time and because they are outside Single User Login, it creates a gulf between those involved in a chapter and those who aren't involved but might have got involved in a specific discussion. I have subscribed to discussions on many talkpages on several wikis, and I get notifications from an eclectic variety of Wikis from among the thousand that the WMF supports. Extending SUL to the wikis of affiliates, or consolidating such wikis with meta, would in my view reduce some of the gulf that I see developing between affiliates and the broader movement. Not sure if that addresses your original concern, and yes I can see that chapters with a right to left script might not want to migrate their wiki to meta. But extending SUL to chapter wikis would at least make it easier to include people in fragmented discussions. ϢereSpielChequers 17:16, 11 September 2026 (UTC)Reply

New update at the WMF Fundraising Hub

[edit]

Hi everyone, We just posted a Q1 update over at the Fundraising Hub. If you are interested to see the work that has been going on and engage with the campaign, go across and have a look. Best, JBrungs (WMF) (talk) 05:01, 10 September 2026 (UTC)Reply

Source Verification Suggestion

[edit]

Hi y'all – one outcome of the recent AI-generated edit suggestions discussions was reinforcing the need for us to be in touch with you all early in the process of exploring new inference-based suggestions. This way, we can discuss the experimental suggestion's risks and decide whether en.wiki might be a good place to evaluate its reliability before considering a wider deployment.

We're now at this point with a new experimental suggestion that we'd value your feedback on.

Inspired by the volunteer-authored AI Source Verifier, and how helpful many experienced editors have found it to be in identifying cases where a citation might not support the claim it's attached to, the Foundation's Research, Machine Learning, and Editing teams are working with Alaexis to try integrating this script as a "suggestion", only visible to experienced volunteers who have opted into seeing experimental suggestions in Suggestion Mode.

Below, you will find more information about how this proof of concept will work and the input we are needing from you all. Before that, a note on why we're prioritizing work on this suggestion right now…

We are prioritizing this exploratory source verification work in response to hearing from volunteers:

  1. How tedious it can be to find potentially unverified claims within an article (e.g. read the claim, click on each source, ensure you can access each source, etc.)
  2. How source verification work is becoming even more important and prevalent as AI increases both A) the ease with which people can add new content to Wikipedia and B) the risk that said content is not supported by the sources it's accompanied by (i.e. contains hallucinated references).

While experienced editors will remain responsible for evaluating whether a source verifies its associated claim, we are seeking to learn whether a tool like this could make the mechanical parts of this wiki work less toilsome.

And in case you're curious, we did a bit of digging to put some numbers to all of this:

  • English Wikipedia has more than 70 million citations.[1]
  • In one study, annotators worked through a few hundred claims and the web pages cited for them, and judged 12% not supported by the source at all.[2]
  • A separate study found that at least 3.3% of facts on English Wikipedia contradict another fact elsewhere on the project.[3]

Note: Neither of those sets represents the encyclopedia as a whole, so the real number is not currently known to us.

How it works

This experimental suggestion will:

  1. Identify all inline URL-based citation(s) in the article
  2. Extract the text that precedes said citation(s)
  3. Retrieve the content of each cited source (and indicate if it's unsuccessful in doing so)
  4. Use an open weight Qwen3.6-27B model, hosted on Wikimedia's LiftWing infrastructure, to compare the claim against the source's contents
  5. Flag claims that the model has deemed to be only partially supported, not supported by omission, or not supported by contradiction.
    1. Note: The model will accompany each conclusion with the passage(s) from the source it's based on, except where the conclusion is "not supported by omission"

Note: This suggestion would not make edits to Wikipedia and can be configured, like other Edit Checks and Suggestions, to show/not show based on a variety of conditions.

We need your input

We will soon be ready to generate an initial experimental dataset that volunteers can use to offer feedback about the usefulness and reliability of this suggestion. First, we need to decide which wikis/languages to include in this dataset.

This leads us to wonder:

  1. Would any of you all be interested in evaluating a batch of these experimental suggestions for en.wiki articles? Note: The suggestions would be made available to you all in a spreadsheet and as a suggestion within Suggestion Mode, visible only to volunteers who have published ≥100 edits and have opted into experimental suggestions.
  2. If so, what types of articles do you think would be helpful for us to include within this dataset? E.g. new articles, articles of a certain quality, etc.

For anyone interested in seeing a demo and talking about this in a voice call, we will be hosting a meeting in the Wikimedia Community Discord on 14 Sep 2026 from 17:00 - 18:00 UTC. We'll of course be responsive here as well.

In the meantime, you can get a sense for how the suggestion works by installing the user script that User:Alaexis and User:LuisVilla have been maintaining.

References

  1. Mario Morvan, "Citation Location Needed", 17 May 2026; measured from 20,000 randomly sampled articles across dated dumps.
  2. Kamoi, R.; Goyal, T.; Rodriguez, J.; Durrett, G. WiCE: Real-World Entailment for Claims in Wikipedia. EMNLP 2023. pp. 7561–7583.
  3. Semnani, Sina J.; Burapacheep, Jirayu; Khatua, Arpandeep; Atchariyachanvanit, Thanawan; Wang, Zheng; Lam, Monica S. Detecting Corpus-Level Knowledge Inconsistencies in Wikipedia with Large Language Models. EMNLP 2025. arXiv:2509.23233.

FAQ

[edit]

Based on some of the questions that volunteers raised in the previous discussion about LLM-backed suggestions.

Community Feedback

[edit]

Thank you for reading and thinking critically about this work! PPelberg (WMF) (talk) 19:09, 10 September 2026 (UTC)Reply

1. Yes. 2. Is it possible to grab the datasets from those two studies (or any other similar) and cross-reference the pages by article-page categories, talk-page categories, and/or talk-page templates (like WP:CTOP templates), to see if any categories or CTOPs are more error-prone than others? That might help prioritize review of any identified problem areas. I would think WP:BLPs, which are all in Category:Living people, would be a top priority for source verification. Also: articles that are new, haven't been edited for a long time, or are tagged with relevant maintenance templates, like {{dubious}} or {{failed verification}}. Levivich (talk) 19:24, 10 September 2026 (UTC)Reply
Is it possible to grab the datasets from those two studies (or any other similar) and cross-reference the pages by article-page categories, talk-page categories, and/or talk-page templates (like WP:CTOP templates), to see if any categories or CTOPs are more error-prone than others? That might help prioritize review of any identified problem areas.
Oh, I think this is a great question/idea. @Alaexis: do you know if we're able to access the datasets used in the studies we referenced so that we could do the comparison @Levivich is describing?
I would think WP:BLPs, which are all in Category:Living people, would be a top priority for source verification.
This intuitively makes sense to me and to be doubly sure: to what extent (if any) would it be accurate for us to understand you suggesting this because of the following reasons?
  1. Like articles within WP:CTOP, the suggestion performing poorly on BLPs is more consequential relative to articles in other categories like Category:Railway lines
  2. Biographies of living people tend to use news, and other web-accessible sources. This means this suggestion should, in theory, be able to retrieve a larger share of the citations used within them.
Also: articles that are new, haven't been edited for a long time...
Good call. This makes sense to me.
...tagged with relevant maintenance templates, like {{dubious}} or {{failed verification}}
Mmm. In essence, you're saying articles within these two categories offer an existing corpus of claims volunteers have identified as failing verification. Accordingly, running the model against those could help us estimate the model's proficiency. Might I be missing/misinterpreting anything here?
A resulting question that comes to mind as I think about this: might articles tagged with {{dubious}} or {{failed verification}} be more likely to contain offline sources? PPelberg (WMF) (talk) 20:40, 10 September 2026 (UTC)Reply
Some of the datasets are available. In fact Semnani et al have made this analysis themselves Articles in the “history” category exhibit the highest inconsistency rate (17.7%), followed by Everyday Life (16.9%) and Society & Social Sciences (14.3%) (Figure 5). The most common error type in history articles is numerical discrepancy. By contrast, categories requiring precise technical knowledge and quantifiable information—such as Mathematics (5.6%) and Technology (9.4%)—show markedly lower rates.
I like the idea of generating edit suggestions for articles with {{failed verification}} templates for Bayesian reasons - if one such citation has been found it's likely that there are more (aka "if you see one cockroach, there are more" principle). Alaexis¿question? 20:50, 10 September 2026 (UTC)Reply
On why BLP, I actually didn't have either of those reasons in mind, but both are good reasons. I was thinking something similar to #1: content that fails verification (not just poor performance of the tool) is more consequential for BLPs than, eg, railway lines. Of all FVs, those are the ones we should find and fix first.
On why the FV tag, yes, and it could also help us estimate human proficiency :-) I think it'd be useful to know whether the model confirms the FVs or finds that the FVs are actually verified -- either way it'd be useful. But I had in mind what Alaexis mentioned: if there's one FV tag, there are probably more untagged FVs in the article. I'd guess the odds are higher than average (but I don't know that for sure).
For the dubious tag, that's like a "might" or "arguably" fails verification, so it'd be useful to have the tool analyze those for a person to review. A statement tagged dubious needs a verification check.
And yeah, I'd guess dubious and FV tagged content is more likely to have offline sources. I don't know the statistics, but I'd guess offline sources are rare and so it's unlikely to be significantly more?
Cool idea btw and well presented. Thanks to the teams for working on this! Levivich (talk) 21:12, 10 September 2026 (UTC)Reply
This sounds great and I'm glad the WMF is working on it. Regarding "what types of articles", I don't particularly see any reason to restrict articles included in the dataset, beyond "we don't want to scan the entire Wikipedia" - is that the reason? Or something else? In solidarity, asilvering (talk) 19:26, 10 September 2026 (UTC)Reply
If "We don't want to scan the entire Wikipedia" is the or a reason, then not scanning articles tagged as unreferenced (Category:All articles lacking sources) or lacking inline citations (Category:All articles lacking in-text citations) are obvious ones to not include as (assuming the tags are correct, which is a different issue) they cannot contain references that can be validated in this manner (~143k articles total). It's also not worth spending the resources attempting to scan references tagged as (permanently) dead, failed verification, or dubious. Thryduulf (talk) 19:55, 10 September 2026 (UTC)Reply
If "We don't want to scan the entire Wikipedia" is the or a reason, then not scanning articles tagged as unreferenced (Category:All articles lacking sources) or lacking inline citations (Category:All articles lacking in-text citations) are obvious ones to not include as (assuming the tags are correct, which is a different issue) they cannot contain references that can be validated in this manner (~143k articles total).
Great spot, @Thryduulf. Excluding articles in these categories for the reasons you named [i] sounds like a great idea to me unless, of course, there is a consequence here I'm not seeing.
It's also not worth spending the resources attempting to scan references tagged as (permanently) dead, failed verification, or dubious...
With regard to {{failed verification}} and {{dubious}}, can you please say a bit more here? Asked another way: what's prompting you to think it would not be worthwhile to scan articles with those two templates present?
Per what Levivich and I were discussing above, I'd been assuming, perhaps inaccurately, that scanning those articles could provide a helpful baseline to compare the model's proficiency against. Might I be missing something here?
Regarding the "tagged as (permanently) dead" bit specifically, I'm assuming the following. Please let me know what (if anything) I might've missed...
1. I assume you are referring to articles that include Template:Permanent dead link
2. If so, I assume the reason for excluding articles that contain ≥1 of these templates would be because these citations are unlikely to have an archived copy the suggestion could retrieve, making a check on them likely to fail
---
i. We can assume articles within All articles lacking source do not contain citations the model can evaluate claims against PPelberg (WMF) (talk) 21:01, 10 September 2026 (UTC)Reply
re failed verification and dubious, I hadn't thought about training the model. Rather I was just thinking that if a human has already tagged a reference as not supporting the associated text there isn't much benefit in an AI suggesting to a different human that it might not support the associated text (this would be a waste of time and resources). I agree that using them to train the model would be useful.
re permanent dead links, I was thinking on a per-citation not per-article basis - read the tag and skip the associated without spending any resources attempting to verify it in the source.
Your comment about archive templates has sparked another thought though, that this tool could highlight potential problems with archives. Firstly, if the url-status parameter is blank or live then it should attempt to verify using the live link or both links, if it is "dead", "deviated", "usurped" or "unift" then it should only attempt to verify against the archive link.
For every citation that has an archive link the following are possible:
  • Live link and archive are identical, both verify the text (no problem)
  • Live link and archive are identical, neither verify the text (problem, but not with the archive but worth flagging to human)
  • Live link and archive differ, but both verify the text (almost certainly not a problem)
  • Live link and archive differ, live link fails verification archive link passes verification (url-status parameter should be changed to deviated, usurped or unfit, but which probably requires human judgement)
  • Live link and archive differ, live link passes verification but archive link does not (flag this for human attention)
  • Live link and archive differ, both fail verification (include this information when flagging this for human attention)
  • Live link is dead, archive verifies text (url-status should be set to dead, possibly flag to something like user:InternetArchiveBot or some other automated task to make changes more widely).
  • Live link is dead, archive fails verification (change the url-status as above and flag to both bots and humans as other citations to the same source may also be dead and pass verification).
Thryduulf (talk) 22:29, 10 September 2026 (UTC)Reply
re failed verification and dubious...I agree that using them to train the model would be useful.
Wonderful. Thank you for walking out what you had been thinking!
...re permanent dead links, I was thinking on a per-citation not per-article basis - read the tag and skip the associated without spending any resources attempting to verify it in the source.
Ah, I see. In concept, what you're describing seems valuable to me. @Alaexis: do you think instructing the model to "skip over" claims that have the Template:Permanent dead link associated with them would require an update to the prompt?
@Thryduulf: in case you're curious, I'm asking about the prompt above because, for now, we're reluctant to make any changes to it. Rationale: a) the prompt has been performing pretty well, b) adjusting the prompt could affect the output in unexpected ways. For these reasons, we're hoping to keep the prompt as-is for this initial round of evaluation.
Your comment about archive templates has sparked another thought though, that this tool could highlight potential problems with archives. Firstly, if the url-status parameter is blank or live then it should attempt to verify using the live link or both links, if it is "dead", "deviated", "usurped" or "unift" then it should only attempt to verify against the archive link.
Oh, this is an interesting set of cases. @Alaexis two resulting questions for you and/or @Isaac (WMF):
1. Do we know how (if at all) the suggestion will behave in these various archive template cases?
2. More broadly, to what extent (if any) would it be accurate for me to think that both a) the LLM could accommodate nuanced instructions of this sort were we to deem them important and b) implementing these instructions would come in the form of an adjustment to the prompt? PPelberg (WMF) (talk) 00:22, 11 September 2026 (UTC)Reply
@Thryduulf If you're curious about how the code selects links to check, you can see the logic and additional notes in the userscript via the extractHttpUrl function. It was written to prefer internet archive links, and then the original link, and only then some of the other archive sites. Checking multiple URLs is possible without changing model prompts but it would still complicate the pipeline as it would functionally double the number of URLs to scrape and model requests to make so would slow things down a good bit. I'll leave that decision to Alaexis but the way I've been thinking about it: the core goal here is verifying the claim, which thusfar we've attempted to do that by fetching the "best" URL. As you raise, there are a variety of additional checks you could do at the same time with the goal of improving the citation (and therefore general Verifiability). I've also thought about inferring the language of the source URL to help fill in the lang parameters on citation templates. These additional checks/actions would complicate the core claim verification task though so I would almost prefer them to be separate at least from an interface perspective, but I am taking note of them.
I'll attempt to answer re: the permanent dead link suggestion as well. It's likely possible to add a check for them (this would also be with the heuristics for extracting links, not the LLM portion of the flow). My thinking: there aren't a ton of them and if they're a dead link then the flow will quickly+gracefully fail anyways (they would never show a suggestion to the end-user). Adding this sort of language-project-specific logic to the code can also complicate efforts to extend this tooling to other language editions in the future. But I'll leave it up to Alex whether it's worthwhile. Isaac (WMF) (talk) 17:22, 11 September 2026 (UTC)Reply
Unfortunately I don't know js so the link doesn't really aid my understanding (but that's not your problem). I don't have a strong feel for what is possible, my suggestions are all things that would are desirable if they are possible. If there are things it isn't going to do (for any reason) but might discover in the process of what it does do then documenting those things somewhere that other humans and/or bots can deal with would be a good thing. Thryduulf (talk) 18:04, 11 September 2026 (UTC)Reply
Yes and please keep the suggestions coming! I mainly wanted to communicate that even if some of these extension ideas don't end up incorporated into this experimental suggestion, they're not being discarded. Isaac (WMF) (talk) 19:48, 11 September 2026 (UTC)Reply
...beyond "we don't want to scan the entire Wikipedia" - is that the reason? Or something else?
Good question, @Asilvering. What you described is accurate, [i] In addition, it would be ideal if this initial dataset includes the types of articles that:
1. You all can imagine this suggestion being the most useful on
2. Includes cases that you think could be particularly complex/not straightforward so we can learn how/if the model fails here
---
i. The size of this initial batch of suggestions will be limited to ~1,000 articles PPelberg (WMF) (talk) 21:07, 10 September 2026 (UTC)Reply
This seems like a perfectly reasonable use of AI on Wikipedia, since it's not actually generating text. I'm looking forward to seeing how it performs. --Ahecht (TALK
PAGE
)
20:07, 10 September 2026 (UTC)Reply
Yeah, this sounds like a wonderful tool. It would be awesome if this could be fitted with a API front end so other tools could build on top of it. A while ago I built something (https://wikirefs.toolforge.org/show?page_title=Bronx+Grit+Chamber) which parses a page, pulls out all the individual claims, and matches them up with citations. But I did a half-assed job of it and got it to the point where it was good enough for my purposes. And I know if fails badly on some referencing styles. It would be great if all that low-level crud could be done once, properly, correctly, and then everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch. RoySmith (talk) 22:03, 10 September 2026 (UTC)Reply
A couple of other things that would be useful... If you don't find the claim in the cited source, look at the other sources cited in the article. It's not uncommon during editing an article for a properly cited statement to get moved but the citation doesn't move with it. Being able to recover the correct pairing would be valuable. Also, if you can't reach a source's URL, see if you can find some other place that has the same item. For example, I'll often cite a NY Times article via the NYT's own archives, which you may not be able to get to because it's behind a paywall, but you can find the same article in ProQuest or some other aggregator. RoySmith (talk) 22:52, 10 September 2026 (UTC)Reply
Combining source verification with finding the edit which added the source to the article often helps with this, could be a useful extentsion. It also helps when existing text in front of a source is modified without reference to the source. CMD (talk) 00:38, 11 September 2026 (UTC)Reply
I think that you can combine the script with "Who Wrote That" already, but you're that it would be useful to see it at a glance. Alaexis¿question? 11:06, 11 September 2026 (UTC)Reply
@RoySmith, it would be great to integrate with TWL and with the Internet Archive to get access to sources that aren't publicly available. Earwig's Copyvio detector has access to TWL so there is a precedent.
For citation that failed with the "Not Supported - Omission" verdict checking other sources in the article makes sense. It's not necessarily cheap - if you have 30 unsupported citations out of 300 total you'll need to run 30*300=9,000 checks, unless there is some kind of screening (the passages nearest where the claim sits or the sources added in the same edits). I tried building a standalone script (User:Alaexis/CNfirmed) to solve the broader problem of finding reliable sources. It turned out to be harder than I thought though. The biggest problem is *where* to look - generic web search is expensive and the alternatives are brittle. Alaexis¿question? 10:56, 11 September 2026 (UTC)Reply
It would be great if all that low-level crud could be done once, properly, correctly, and then everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch.
@RoySmith to be doubly sure I'm following, by "low-level crud" are you referring to things like splitting the article up into its constituent claims, identifying the source(s) associated with each, etc.?
A while ago I built something...which parses a page, pulls out all the individual claims, and matches them up with citations
Neat! If you happen to have them handy, we'd be eager to learn what referencing styles you noticed this tool struggling with.
...everybody else who wanted to build tools in that space could just take advantage of it instead of reinventing it from scratch.
I can imagine the creativity something like this could inspire. With this said, I think we're still a ways away from being able to determine how feasible something like this would be. Once you confirm the first question I posed above, I'm thinking I can create a phabricator ticket so that we can come back to it at a future point. PPelberg (WMF) (talk) 00:39, 11 September 2026 (UTC)Reply
Yeah, I envision some kind of network API where you give it a page title (revid, whatever) and it gives you back a list of claims and the associated citations in some structured form. Of course, this is really just one specific example of a general pattern. You really should be able to treat every page as an object on which you can perform operations and access those operations via a network API.
We spend too much time and effort reinventing wheels. For example, I don't know how many tools we've got which measure the readable prose size of an article. They all give different results because they all have slightly different definitions of what "readable prose" means. Neither are more right than any other, they're just different. And it's dumb that so many people have spent so much time writing essentially the same function in slightly different ways.
To go back to this particular example, there's already a bunch of different referencing styles in use. My code parses one of them (the one I tend to use) pretty well. I know it fails on others (to be honest, I don't remember which). Imagine a world where parsing claims and citations from articles was handled by one API that handled them all. Then all sorts of tools could take advantage of that. And more to the point when a new format comes along (say, sub-referencing, which is going to hit enwiki Real Soon Now), one bit of code will need to be adapted to handle that, and automatically all the various other tools will now be able to handle it too. RoySmith (talk) 01:12, 11 September 2026 (UTC)Reply
@RoySmith, that's so true. I had to build a lot of non-core stuff and would've been happy to reuse others' work instead. Discoverability of tools is a huge topic, Toolhub-Evolved is the latest solution I'm aware of.
Currently verification is not exposed as a standalone service but that wouldn't be too hard to do if you have a use case in mind. Fetching website data is already a standalone service on Toolforge (repo, ToolHub page). Alaexis¿question? 10:15, 11 September 2026 (UTC)Reply
Thank you for bringing this discussion here. As someone who's used Alaexis's tool on and off for a few months, this is the kind of collaboration I like to see: working with an established editor to support development or rollout of a tool that's already been field-tested by the community.
To answer your questions: 1. Yes, definitely. 2. Dreamyshade has been building https://projo.toolforge.org/review to identify high-priority articles for WP:NPP reviewers – another opportunity for collaboration or sharing ideas? —ClaudineChionh (she/her · talk · email) 00:51, 11 September 2026 (UTC)Reply
Yes, thank you Claudine! The Projo unreviewed article priority-ranking tool (which is new and still in development, still janky) reflects my understanding of how to prioritize articles that need editor help in general, based on factors including CTOPs, BLPs, reference need, page views, orphan status, and cleanup tags such as AI-generated, COI, and POV (the page contains a table of all the factors). It's focused on supporting New Pages patrollers because that is the most urgent work that I know of right now, with the giant backlog and constant influx of COI/UPE and LLM-generated articles. I'm adding more factors based on data I can derive from categories and other elements in the replica database. I'd love to have access to APIs for predicted percentage of LLM-generated content and predicted numbers of verified/partially-verified/failed-verification citations!
I use Alaexis' Source Verifier a lot, especially for Articles for Creation review, New Pages patrol, and AI cleanup tasks. I find it very helpful, especially on articles that are tagged as likely AI-generated or that I suspect are AI-generated. It helps me rapidly check source-text integrity, which helps me figure out whether an article is mostly fine, salvageable, or trash. It's part of my standard toolkit for article quality evaluation, along with https://copyvios.toolforge.org/, https://wikipedia.gptzero.me/, and Cite Unseen - none of them are perfect, just tools, but I believe they're useful when used with competence and a grain of salt. I've been trying to gather lists of tools along these lines: Article workflows#Tools for specific workflow steps + User:Dreamyshade/Article workflows#Semi-automation. (That page has some of the ideas I'm trying to put into practice in the Projo tool.) Dreamyshade (talk) 01:48, 11 September 2026 (UTC)Reply
Wouldn't it not be possible to provide this System access to paywalled content over the Wikipedia Library? The Other Karma (talk) 09:44, 12 September 2026 (UTC)Reply