Mathematical Prompting
yoyleguy <[email protected]>
| Newsgroups | comp.ai |
|---|---|
| Organization | A noiseless patient Spider |
| Message-ID | <[email protected]> |
I remember reading that there was a way to bypass the safety filters of an AI model by posing the prompt as a mathematical problem (e.g. asking it how to rob a bank by involving the removal of security systems as an element of the problem) with high success rates. Why haven't they solved it by simply just killing off the response if it appears at all to be providing instructions for or inciting criminal activities? It's likely few false positives will appear, so I'd say it's worth the risk. -- your local idiot and a young one? yes i am y:g btw