CF tries to help with the judgment and make it easier, but doesn’t fundamentally change it. Binary judgments tend to be easier than non-binary judgments. And CF’s emphasis on evaluating ideas for goals, and putting work into clarifying the goal more than is normal, and breaking the goal into sub-goals that can be evaluated separately, helps.
FYI, CF criticizes MCDM.
CF has some ideas about criticism breadth (designed for humans, so it may not all help agents), like Paths Forward, public intellectuals having written debate policies and being open to some debate with the general public, discussion and debate methodology ideas (including using tree diagrams to organize things and using more literature cites and writing reusable texts when no existing literature covers an issue), and the articles about engaging with intuitions. Those help but it’s still hard.
Also relevant (and found in CR and elsewhere more) is intellectual tolerance, humility, being curious, not being dismissive, and not judging ideas by bad proxies like popularity or the social status of the person who thought of it or the person currently advocating it.
I used to think if I said something to people [particularly Popperians on intellectual discussion forums], and they had evidence that I was wrong, they would tell me.
Your point about intellectual tolerance and humility brings up an interesting AI design question. Would an AGI have an in-built advantage here because it doesn’t naturally have human social and emotional baggage? Or is it that any system capable of judgment must inherently start with some framework of initial heuristics and assumptions?
I’m thinking about how Deutsch interprets Gödel that it really just means you can’t mechanically derive all truth from a fixed, closed set of axioms. You need creativity to guess new explanations and step outside the initial system.
Translating that to agents: does that mean it’s formally impossible to build a ‘complete’, perfectly neutral judging machine that just mechanically outputs good judgments? And therefore an agent must be intrinsically fallible and reliant on its initial guesses, which means absolutely all the pressure to not stay wrong falls on having a robust error-correction environment?
In general, AGIs would be people and be capable of learning from other people, having someone interact with them in the parent role, having teachers, being influenced by culture, and learning social and emotional baggage.
It’s unknown how much physiological differences (different bodies missing human hormones and other chemicals) may make some memes not apply and sort of break that knowledge for AGIs which are in a different context. It’s hard to say which memes depend how much on human chemistry. I’d very loosely guess social status has under 10% dependence and emotions under 50%. I disagree with people who think ~all emotions depend on human chemistry (or on alternative emotional systems that AGIs wouldn’t have by default – that would have to be intentionally designed and added).
Fallible agents that create knowledge like humans (using C&R aka evolution of ideas) is what I would expect. I don’t think we know an alternative approach to AGI. It’s harder to comment on what’s impossible.
one of the most fascinating ideas I came across was on sam harris podcast it would take a lot of work to understand it deeply. deutsch said something like epistemology is substrate independent. what does that mean? it sounds like it contradicts what you say here
Substrate is the thing it’s made out of. It’s more typically discussed with computers: computation has the same rules whether you build a computer out of silicon or any other materials.
Deutsch is saying that epistemology would work the same for AGIs even though they are made of different materials than humans are made of. He also believes that different AGI designs won’t change epistemology, similar to how different (classical, non-quantum) computer designs are unable to change what algorithms a computer can run (unless you handicap the computer – you can get it to run less but not more).
oh ok thanks. but doesn’t that imply the details about human chemistry will fall away? if epistemology is substrate independent and AGIs are on silicon, human chemistry won’t play a role in their thinking. but they’d still face the same epistemological challenges of being fallible, needing criticism, depending on the environment to catch errors?
Yes. Although you could emulate the effects human chemistry have on humans if you wanted to even if it isn’t necessary to achieve intelligence. I think qualia isn’t a direct consequence of chemistry, but rather something the mind creates in response to certain thoughts and(or?) sense data. Qualia is one point Elliot and Deutsch have disagreed on since a long time ago (Curiosity – I Changed My Mind About David Deutsch), although I don’t know what their positions are on it. I would be interested to know @Elliot.
I think most discussion of qualia is unproductive and that the concept is unnecessary for most or all philosophy. I think the concept caught on for its ability to confuse and awe people, not its usefulness. Deutsch likes talking about qualia.
going back to the AI angle. in your epistemology article you listed the hard subproblems for AI:
represent ideas in code
represent criticism in code (implied by 1)
detect which ideas contradict each other
brainstorm new ideas and variants
and then the really hard problem: when two ideas contradict, which one is wrong?
it seems like LLMs already do a decent job at 1, 2, and 4. they can represent and generate ideas, generate criticisms, and brainstorm variants. 3 is partial, they can often spot contradictions but not reliably. the part they can’t do is the judgment step we were discussing earlier, evaluating which idea to reject when they contradict.
does that mapping look right to you? and if so, does that change your view on how close current systems are, like they’ve accidentally solved the easier parts and the remaining bottleneck is exactly the judgment problem you identified?
actually one more data point on the mapping. the claude mythos results look a lot like step 3 (detecting contradictions) working at a superhuman level, at least in code. finding a vulnerability is basically detecting a contradiction between what the code is supposed to do and what it actually does.
does that change the picture at all? or would you say code contradictions are not the same as contradictions between ideas/explanations?
I don’t think LLMs represent ideas in code at all. They have tokens, not ideas. There is no idea data structure, and no function to apply an idea as a criticism of another idea.
i think these models can spot contradictions in some of your philosophical articles as well that a human reader would agree look like real contradictions. not necessarily decisive ones, but ones worth addressing. wouldn’t that be step 3 working in the ideas domain too?
wouldn’t it be fair to apply the substrate independence concept here or would it be a category error? if i give a model one of your articles and it outputs “paragraph 4 says X but paragraph 12 says Y, and those contradict because Z” and a human reader checks and agrees that’s a real contradiction, does it matter that the model used tokens instead of an idea data structure to find it? we’re going by does anyone see an error and someone saw the error.
A difference is LLMs can basically only find the types of contradictions that many people have found and written down before. They rely on training data and existing patterns. They don’t understand things conceptually which limits their ability to deal with new ideas.
I can grant that but even then if the models were only catching things that have already been figured out by someone else, wouldn’t that still be incredibly helpful for error correction? it would still act as an automated check to tell you if you’re making an already-known error.
i experiment a lot with harness and i kind of have one for this as well. if you’re interested i can share some examples
When humans use words, the words correspond to ideas in their brain. When LLMs output words, that correspondence is missing. We’re used to inferring that a human saw an error (had ideas about the error) when they say certain words (how else would they produce those words?), but this doesn’t apply in the same way to LLMs.
LLMs don’t do error correction on the words they output. I don’t think you should consider an error to be found if the claim hasn’t yet received critical thinking to check for errors.
LLMs pattern match training data that already had error correction applied to it by people. This makes the outputs look kind of like they’ve had their own error correction even though they haven’t. Humans can’t output words that well without error correcting them, so we’re used to seeing words that form coherent sentences and inferring they were error corrected.
LLMs are sort of like a very fast literature review: if people already wrote down something relevant, it may find it. That’s different than LLMs doing their own thinking.
(As always, this is my current understanding. These are conjectures for discussion.)
1. the essay’s policy claim is that government should police companies more vigorously. but the explanatory claim is that the root problem is human irrationality. you write:
“The people who run big companies are similar to the people who run the government. They aren’t significantly better people.”
“Most people are pretty bad. The main types of badness I have in mind are irrationality and dishonesty.”
if the government is staffed by the same irrational, dishonest people who run the companies, what makes the government capable of policing them effectively? the essay says the fix is education via your philosophy work, but that means the policy proposal has a precondition that doesn’t exist yet. is the harness missing a mechanism you had in mind, or is the actual claim “more policing would be good if we first fixed the people problem”?
2. you argue that creative adversaries with enough resources will generally find a way to beat any fixed set of rules. and then you say the solution requires “creative judges and law-makers on an ongoing basis” because companies will find loopholes. but a government that must constantly adapt its rules to outmaneuver well-resourced corporate adversaries is running an open-ended regulatory arms race. you call this a minarchy earlier in the essay, but a state that needs to constantly outmaneuver billionaire-funded corporate adversaries doesn’t seem minimal. your own premise (creative adversaries beat fixed rules) seems to undermine your own proposed solution (minarchist rules enforced better).
did the models hallucinate these, or are they real contradictions you’ve thought about since writing it?