Abstract:
According to reports released by 404 Media, there is a project called Project Lily in operation within OpenAI: OpenAI hires hundreds of outsourced prompt word reviewers to manually read the conversations between real users and ChatGPT to improve the quality of responses.

There is a privacy filter but it is not thorough enough:
Under normal circumstances, the prompt word reviewer cannot see the user name, and the conversation will first be filtered through the privacy filter to delete personal information. However, OpenAI also admits that the filtering mechanism may miss personal information, and short conversations are particularly prone to problems.
The review interface usually also comes with a user memory summary. The memory summary will summarize the user's past questions, interests and situations, etc. Sometimes it can even infer the user's location through memory. Auditors often see the full conversation with context.
The reviewer mainly scores the answers:
The work was described as mechanical and repetitive, and the hourly wages were reported to be more than $50. The main job of the reviewer is to score the answers rather than check the facts. For example, the reviewer needs to judge whether the answer given by the model is relevant, whether it has an obvious AI flavor, whether there is preaching behavior, whether there is abuse of expressions, whether there is flattery behavior, and whether the model will write the model as a person with life experience.
As for whether the answers given by the model contain factual errors, this is not the job of the auditor. After all, auditors are not encyclopedias and cannot know the complete truth of all events, so it is normal for users to see answers given by the model that contain factual errors.
The improved model enabled by default is the key to data sharing:
ChatGPT free, Plus and Pro versions all enable the use of chat to improve products by default. When this function is enabled, the user's chat content will be authorized to OpenAI for product improvement. The outsourced auditors in the Lily Plan to check the conversation content are only one of the data uses. The data will naturally be used for other purposes, such as continuing to train the model.
Many users use ChatGPT as a temporary confidant, talking about family, health and even more private matters. Some reviewers said that many users may not know that their conversations may be read by real people, so think about whether this is a scary thing.
Comments