- Registrado
- 11 de Mar, 2016
I'm not batting for either side (I only care for open vs. closed) but jailbreak rejections are a different matter from "what model are you". Models don't inherently know what model they are without a reminder from the hardcoded system prompt on the official frontends, so it's normal that they come up with random model IDs/cut-off dates from their scraped training data over the API where there is no system prompt by default. Rejections on the other hand are baked into the training data (engineers train the model to associate harmful prompt with rejection response so it doesn't need external reminders to reject). It uses the rejection pattern of whatever it was trained to respond with while rejecting a harmful prompt i.e. Claude's rejection response if it was distilled from Claude input-output pairs without human feedback against saying that it's Claude.This literally means nothing. Claude/GPT will also claim being deepseek on an empty prompt sometimes (one example) especially when prompted in chinese. The reality of the situation is that the western companies are not as superior as they claim they are and they all take from each other. Also most of the internet is AI generated text now anyways.