I found this was a bigger problem with the older models, the newer models are much more "neutral", if you exclude their system prompts they get on the websites, of course. Personality is a question of prompting nowadays, the strength of in-context learning is downright uncompareable to what was only 1-2 years ago. Slop, be it in image generation or text, is almost always a skill issue when generating AI content.
For a simple example - everyone knows GPTisms. "It's not x, it's y", glazing the user, my newest favorite the "gold standard" etc.. Now, the intuitive way to solve this is simply telling the model to stop doing it. Everyone who ever did anything with an LLM that goes past chatting with it on a website probably tried this at some point. Doesn't really work reliably. The biases are too strong. You can try fighting it mathematically with samplers (even though I feel people gave up on this) but the results are a bit hit and miss. What people do not try, (because it doesn't feel intuitive) is for example to feed the last output of the LLM again into the LLM (with no context) and clear instructions to fix these issues in the given text. You can even do multiple passes to scan for different issues if you like. You could tell the LLM in the second pass to structure the incoming text like an example text you provide. You can also work some regex magic by removing specific structures from the given text (leaving a word salad) and have the LLM reconstruct it, which funnily in my experiments often will change the tone of the text fundamentally, but the content a lot less than you think. Voila, you have your AI text that doesn't really sound like an AI text anymore. In my experience people don't even attempt solutions like this simply because they don't feel like "how this is supposed to work".
I don't see these biases as a failing or weakness of the technology btw. as even humans, the actual gold standard if you will, aren't really any different in that regard.
Alignment will always stay impossible, as these models all lack ground truth. When we get to AGI, I think it'll be a bit more complicated than "LLM".
---
Kimi K3 was released today and from benchmarks, it's trading blows with claude opus/fable and gpt 5.6, while costing as much as sonnet. *If* true, that is going to be interesting and would mean that China has officially caught up. Almost more importantly, it's gonna be open weight. (even though at I think 2.7T, you're not going to run it) Anthropic/OpenAI whining about China stealing tech/warning about dangerous AI!!!11! incoming in 3...2...1...