If you're not a contest judge and just want agent skills for ingesting and improving arguments, go to README.md.
This is about a 10-12 minute read.
Judging this contest should be an easy job, and almost is
The FLF has asked for tools and methodologies to make reasoning easy to scrutinize. Assume a contestant has made such a tool or methodology. The contestant should be able to use their tool or technology to lay out reasoning for why their submission should win. It's then easy to scrutinize that reasoning, and thus easy to see whether or not their submission should win.
The caveat is, you've never used this tool or methodology before. If such tools and methodologies were common already, why make a contest to create one? They're either nonexistent or uncommon, so even if the tool or methodology is merely a new combination of existing concepts, there's going to be a learning curve.
This is not hypothetical. The submission you're judging now is a methodology and tool for making reasoning easy to scrutinize. The next sections will walk you through the learning curve, and then you'll scrutinize my reasoning for why this submission is a winner.
What "easy to scrutinize" looks like: HowTruthful
Spoiler alert: The tool is called HowTruthful. Rather than give step-by-step instructions for using it, I'm going to explain the reasoning and motivation behind how it works. Then how to use it will click right away.
The obvious way to represent reasoning that everyboy missed
For most of the exactly-200-year history of argument maps, they've been made using ink or pencil on paper. One innovation from their original form was to draw circles around the statements so that they don't run into each other on the paper. Another was to draw directional arrows instead of symmetric lines, so that conclusion-to-premise could be drawn in any direction, not just downward on the paper. Finally, we got computers. There was no longer any edge to the paper, and the circles could be moved to make more room whenever a new one came in.
Arrows connecting circles in two dimensions. That's argument maps since, at the latest, 1958. And when software engineers see arrows connecting circles in two dimensions, they recognize a graph. Software engineers should know that there are other ways to represent graphs besides two-dimensional circle/arrow diagrams. The most prominent example is a hypertext web, ubiquitous to the point where "Internet" and "web" are often used interchangeably.
Somehow, the idea of using hypertext to represent the graph of an argument map is so invisible that even Scott Alexander, a knowledgeable and insightful blogger prominent in the rationalist community, when writing about the abundance of argument-mapping projects, writes as if the circles-and-arrows representation is the only one. "Once you have enough of these circles, aren’t you fighting the argument-mapping idea rather than benefiting from it?" Similar objections are noted on Wikipedia.
When you put statements in circles and connect them with arrows in two dimensions, you run into scaling problems with large numbers of statements. When you put statements in pages and connect them by hypertext, you scale much better. Every statement has a page where you look primarily at the statement, and secondarily at its immediate pro and con connected statements. What you're looking at is essentially a high-level summary. You click into a pro or con statement to dig deeper. It scales to however many statements you want.
Scrutiny and assessing truthfulness are intertwined
Picture yourself looking at the highly-focused format described in the previous section: a statement, and a high-level summary of why you should or shouldn't believe it. Why are you looking at it? You're looking at it in order to decide how truthful it is. Why else would you scrutinize it?
Every statement on HowTruthful is accompanied by a colorful 1-5 rating scale. Everything starts out as a colorless 3, debatable. When there are debatable pros and cons, you click into them, until you reach a statement that's self-evidently true or false, or that has enough non-debatable pros and cons for you to decide its truth. Then a single click changes the colorless 3 into one of the colorful truth values. The process of navigating down through the argument map, and adding color on your way back up, is fun.
For this reason, I've made no attempt in this submission to automate the assessment step with AI. If you really want to let an AI assess truthfulness in a file you want to import to HowTruthful, you can probably just ask it. I haven't tried, though, because the whole point of letting a human scrutinize is to let a human assess.
Where AI proves useful
Clicking the pretty colored rating discs is the fun part of using HowTruthful. The tedious part is creating the graph of statements. You type in "The sky is blue." You click through to its page and stare at it. You decide you need some evidence before you can rate its truthfulness. You click the "Pro" header and type in "It looks blue." Then you click through to that statement's page. You notice that this statement is context-dependent and click it, and edit to "The sky looks blue." You click the Save button and continue.
We have computers. Computers process information. Why not have the computer process the freeform text you were looking at when you decided you wanted to scrutinize reasoning, and transform it into a web of context-independent statements linked by pro and con relationships? If you had asked me this question before modern LLMs came out, I would have laughed and told you computers don't work that way. But today I'd answer that that's a great idea.
Why not integrate AI directly into the HowTruthful web interface?
My vision for HowTruthful is a place where adversaries can meld their arguments and arrive at what, for them, are cruxes. It needs to be a platform people trust. Having a single built-in AI for making the initial draft of an argument would rightly lead people to wonder if bias was secretly being introduced. For this reason, I think it's important to let people drive the AI parts of the process from their own choice of agent, using skills that they can inspect and modify themselves.
That's it for backround and motivation. Now it's time to try it.
Options for trying it out
Large
Use Claude Code or your favorite alternative to open this repo as a project. Follow the README.md instructions to install optional prerequisites and start prompting. It may take several minutes for your LLM to ingest a large corpus. Try pointing it at the contest announcement and asking it to ingest the links for the 3 case studies.
Ask it to import what you ingested. You'll be taken to a HowTruthful page where you scroll down and click Import. Then start clicking statements as described above.
Medium
If you trust me, you can skip trying out the LLM agent skills yourself and just believe my descriptions of how I used the skills to create the examples below. Then do the click / scroll down / import thing to assess how well what I did worked.
With the eggs case study, I used an early version of the skill. It produced an imperfect graph. I fixed up one branch of the argument and rated truthfulness according to my own opinions looking at the evidence.
After vibe coding the howtruthful-ingestion skill for another week, I mapped the LHC argument. Click The LHC will not create a black hole that destroys the Earth, scroll down, hit the Import button, and explore. Click those colored discs to rate truthfulness. This will only be stored in your own web browser on your own device.
With the COVID-19 origins case study, I spent significant time vibing the argument together with claude. This is where a lot of refinements of the howtruthful-ingestion skill came from. The argument is large and will take a while to explore.
This case study illustrates HowTruthful's ability to let work build on work. Import the judges' decisions file first. Same process as above: click the link, scroll down, press the Import button, see the new statement among "Your Private Opinions" top-level statements.
That's going to change when you use the import link below. It's going to find your existing opinions and graft them in wherever the same statement appears. You're constructing one big argument out of several small ones.
When manually putting arguments in on HowTruthful, autocomplete suggests existing opinions. If you choose one, or even if you accidentally type in the same thing you typed before, it will reuse that opinion. In this way it becomes a collection of reusable knowledge.
Small
Now, for the delightlfully self-referential part I promised at the beginning. As a judge, you're asking, "Is Epistack-HowTruthful" a winner. As with any question, you can transform it into "How truthful is the statement, 'Epistack-HowTruthful is a winner.'?" I already put that statement in HowTruthful. Then I added statements under the Pro section that, if true, would lend truth to the main statement. I copied those from your judging criteria. In turn, each of those statements got sub-statements copied from the same doc that, if true, would lend truth to them. Unlike prose, where we assume a default of "true", these all started out at the default rating of "debatable". This is a perfect starting point for answering complex questions like the one you're judging now.
I navigated through this tree of statements. A few were self-evident, where I just changed the rating without adding any Pro or Con underneath them. For many I did add pros and cons, debating each statement before rating its truthfulness. I may be biased about this submission being a winner, but at least I know my reasoning is clear.
As promised, you can go to the import page prefilled with my reasoning, scroll down, click Import, and navigate through it. When you see something you want to verify for yourself, change it to 3 - debatable. Investigate everything, put in your rating for that statement, then walk back up the tree, reassessing based on all the evidence. At the top, you might have a clear picture as to whether this submission is a winner.
Or you might not have a clear picture. I deliberately left the last part of the judging criteria debatable because I felt it would be presumptuous to rate that with my own opinion. That is one of Claude's many criticisms of me underselling this submission. Judge for yourself.
Future work
As I mentioned in an earlier section, I think it's important to let people drive the AI parts of the process from their own choice of agent, using skills that they can inspect and modify themselves. For this reason I don't have grand plans to make AI part of everybody's UI. But there are still useful things I can do.
The skills in this repo were hastily thrown together, with no hesitation to rely on context that AGENTS.md provides. Skills are supposed to work when copied into other places, so I need to clean up where the context lives. There are also decisions to make about how many skills HowTruthful-LLM integration should comprise.
There's a premium option on HowTruthful where opinions are stored in a cloud database and can be shared with everyone. An MCP server would be a helpful addition to premium, letting people use their own choice of agent to help them construct complex arguments to assess.
The premium option needs a lot of work. It's weak in the import/export area. And it needs to be much more social. I have ideas for making it such that one can find contrasting opinions from any opinion page. I imagine people of diverse opinions coming together to find common ground and understand each other better. It might not be as hard as commonly thought. I'm going to try.
Closing thoughts
I watched Twitter succeed when its main distinguishing characteristic was that it limited posts to 140 characters. This totally non-novel, small design decision had surprise effects by creating a platform where everything was concise. I copied Twitter in this respect, but of course, as a child of the 8-bit computing area, I used the one true length limit of 255.
HowTruthful also has its own small design decisions that aren't particularly novel technologically, but have nice surprise effects. The default value of "debatable" as you enter new statements starts to feel really nice when you're brainstorming, for example.
But perhaps the most surprising was when I compromised on the "make everything as clean and minimialistic as possible" principle I mostly follow, and removed the code that hid the "Con" section header when it was empty. Actually it was both "Pro" and "Con", but the nice surprise effect was all about "Con".
Leaving the "Con" header floating there over empty space turned out to be a reminder to always consider contrasting opinions. So I'll close doing that.
My opinion is that there's tremendous value in the dispassionate reasoning facilitated by structuring arguments the way HowTruthful does. A contrasting opinion is that there's value in passionate prose. It's always entertaining when passionate prose emerges unexpectedly from an LLM. Of course, the LLM itself has no passion. It doesn't even have an underlying meaning to its words. Any "meaning" in its model is just relative to other opaque words. But the prose that pops out is a statistical prediction based on everyone whose thoughts went into its training data.
So when I was doing my due diligence and having Claude review the argument map I'd made about whether this submission was a winner, all those people whose thoughts went into the training data took me to task severely for underselling it. I didn't want to translate their points into a graph of dispassionate context-independent statements. I can't believe I'm saying this, but please read this AI-generated text: claude-assessment.md
The contest deadline was today. Here's what I submitted. Original here.
Contest submission: Epistack-HowTruthful
This is a submission to FLF's Epistemic Case Study Competition.
If you're not a contest judge and just want agent skills for ingesting and improving arguments, go to README.md.
This is about a 10-12 minute read.
Judging this contest should be an easy job, and almost is
The FLF has asked for tools and methodologies to make reasoning easy to scrutinize. Assume a contestant has made such a tool or methodology. The contestant should be able to use their tool or technology to lay out reasoning for why their submission should win. It's then easy to scrutinize that reasoning, and thus easy to see whether or not their submission should win.
The caveat is, you've never used this tool or methodology before. If such tools and methodologies were common already, why make a contest to create one? They're either nonexistent or uncommon, so even if the tool or methodology is merely a new combination of existing concepts, there's going to be a learning curve.
This is not hypothetical. The submission you're judging now is a methodology and tool for making reasoning easy to scrutinize. The next sections will walk you through the learning curve, and then you'll scrutinize my reasoning for why this submission is a winner.
What "easy to scrutinize" looks like: HowTruthful
Spoiler alert: The tool is called HowTruthful. Rather than give step-by-step instructions for using it, I'm going to explain the reasoning and motivation behind how it works. Then how to use it will click right away.
The obvious way to represent reasoning that everyboy missed
For most of the exactly-200-year history of argument maps, they've been made using ink or pencil on paper. One innovation from their original form was to draw circles around the statements so that they don't run into each other on the paper. Another was to draw directional arrows instead of symmetric lines, so that conclusion-to-premise could be drawn in any direction, not just downward on the paper. Finally, we got computers. There was no longer any edge to the paper, and the circles could be moved to make more room whenever a new one came in.
Arrows connecting circles in two dimensions. That's argument maps since, at the latest, 1958. And when software engineers see arrows connecting circles in two dimensions, they recognize a graph. Software engineers should know that there are other ways to represent graphs besides two-dimensional circle/arrow diagrams. The most prominent example is a hypertext web, ubiquitous to the point where "Internet" and "web" are often used interchangeably.
Somehow, the idea of using hypertext to represent the graph of an argument map is so invisible that even Scott Alexander, a knowledgeable and insightful blogger prominent in the rationalist community, when writing about the abundance of argument-mapping projects, writes as if the circles-and-arrows representation is the only one. "Once you have enough of these circles, aren’t you fighting the argument-mapping idea rather than benefiting from it?" Similar objections are noted on Wikipedia.
When you put statements in circles and connect them with arrows in two dimensions, you run into scaling problems with large numbers of statements. When you put statements in pages and connect them by hypertext, you scale much better. Every statement has a page where you look primarily at the statement, and secondarily at its immediate pro and con connected statements. What you're looking at is essentially a high-level summary. You click into a pro or con statement to dig deeper. It scales to however many statements you want.
Scrutiny and assessing truthfulness are intertwined
Picture yourself looking at the highly-focused format described in the previous section: a statement, and a high-level summary of why you should or shouldn't believe it. Why are you looking at it? You're looking at it in order to decide how truthful it is. Why else would you scrutinize it?
Every statement on HowTruthful is accompanied by a colorful 1-5 rating scale. Everything starts out as a colorless 3, debatable. When there are debatable pros and cons, you click into them, until you reach a statement that's self-evidently true or false, or that has enough non-debatable pros and cons for you to decide its truth. Then a single click changes the colorless 3 into one of the colorful truth values. The process of navigating down through the argument map, and adding color on your way back up, is fun.
For this reason, I've made no attempt in this submission to automate the assessment step with AI. If you really want to let an AI assess truthfulness in a file you want to import to HowTruthful, you can probably just ask it. I haven't tried, though, because the whole point of letting a human scrutinize is to let a human assess.
Where AI proves useful
Clicking the pretty colored rating discs is the fun part of using HowTruthful. The tedious part is creating the graph of statements. You type in "The sky is blue." You click through to its page and stare at it. You decide you need some evidence before you can rate its truthfulness. You click the "Pro" header and type in "It looks blue." Then you click through to that statement's page. You notice that this statement is context-dependent and click it, and edit to "The sky looks blue." You click the Save button and continue.
We have computers. Computers process information. Why not have the computer process the freeform text you were looking at when you decided you wanted to scrutinize reasoning, and transform it into a web of context-independent statements linked by pro and con relationships? If you had asked me this question before modern LLMs came out, I would have laughed and told you computers don't work that way. But today I'd answer that that's a great idea.
Why not integrate AI directly into the HowTruthful web interface?
My vision for HowTruthful is a place where adversaries can meld their arguments and arrive at what, for them, are cruxes. It needs to be a platform people trust. Having a single built-in AI for making the initial draft of an argument would rightly lead people to wonder if bias was secretly being introduced. For this reason, I think it's important to let people drive the AI parts of the process from their own choice of agent, using skills that they can inspect and modify themselves.
That's it for backround and motivation. Now it's time to try it.
Options for trying it out
Large
Use Claude Code or your favorite alternative to open this repo as a project. Follow the README.md instructions to install optional prerequisites and start prompting. It may take several minutes for your LLM to ingest a large corpus. Try pointing it at the contest announcement and asking it to ingest the links for the 3 case studies.
Ask it to import what you ingested. You'll be taken to a HowTruthful page where you scroll down and click Import. Then start clicking statements as described above.
Medium
If you trust me, you can skip trying out the LLM agent skills yourself and just believe my descriptions of how I used the skills to create the examples below. Then do the click / scroll down / import thing to assess how well what I did worked.
With the eggs case study, I used an early version of the skill. It produced an imperfect graph. I fixed up one branch of the argument and rated truthfulness according to my own opinions looking at the evidence.
Click Are eggs good to eat?, scroll down, hit the Import button, and explore.
After vibe coding the howtruthful-ingestion skill for another week, I mapped the LHC argument. Click The LHC will not create a black hole that destroys the Earth, scroll down, hit the Import button, and explore. Click those colored discs to rate truthfulness. This will only be stored in your own web browser on your own device.
With the COVID-19 origins case study, I spent significant time vibing the argument together with claude. This is where a lot of refinements of the howtruthful-ingestion skill came from. The argument is large and will take a while to explore.
This case study illustrates HowTruthful's ability to let work build on work. Import the judges' decisions file first. Same process as above: click the link, scroll down, press the Import button, see the new statement among "Your Private Opinions" top-level statements.
Now notice that you have a new top-level opinion.
That's going to change when you use the import link below. It's going to find your existing opinions and graft them in wherever the same statement appears. You're constructing one big argument out of several small ones.
When manually putting arguments in on HowTruthful, autocomplete suggests existing opinions. If you choose one, or even if you accidentally type in the same thing you typed before, it will reuse that opinion. In this way it becomes a collection of reusable knowledge.
Small
Now, for the delightlfully self-referential part I promised at the beginning. As a judge, you're asking, "Is Epistack-HowTruthful" a winner. As with any question, you can transform it into "How truthful is the statement, 'Epistack-HowTruthful is a winner.'?" I already put that statement in HowTruthful. Then I added statements under the Pro section that, if true, would lend truth to the main statement. I copied those from your judging criteria. In turn, each of those statements got sub-statements copied from the same doc that, if true, would lend truth to them. Unlike prose, where we assume a default of "true", these all started out at the default rating of "debatable". This is a perfect starting point for answering complex questions like the one you're judging now.
I navigated through this tree of statements. A few were self-evident, where I just changed the rating without adding any Pro or Con underneath them. For many I did add pros and cons, debating each statement before rating its truthfulness. I may be biased about this submission being a winner, but at least I know my reasoning is clear.
As promised, you can go to the import page prefilled with my reasoning, scroll down, click Import, and navigate through it. When you see something you want to verify for yourself, change it to 3 - debatable. Investigate everything, put in your rating for that statement, then walk back up the tree, reassessing based on all the evidence. At the top, you might have a clear picture as to whether this submission is a winner.
Or you might not have a clear picture. I deliberately left the last part of the judging criteria debatable because I felt it would be presumptuous to rate that with my own opinion. That is one of Claude's many criticisms of me underselling this submission. Judge for yourself.
Future work
As I mentioned in an earlier section, I think it's important to let people drive the AI parts of the process from their own choice of agent, using skills that they can inspect and modify themselves. For this reason I don't have grand plans to make AI part of everybody's UI. But there are still useful things I can do.
The skills in this repo were hastily thrown together, with no hesitation to rely on context that AGENTS.md provides. Skills are supposed to work when copied into other places, so I need to clean up where the context lives. There are also decisions to make about how many skills HowTruthful-LLM integration should comprise.
There's a premium option on HowTruthful where opinions are stored in a cloud database and can be shared with everyone. An MCP server would be a helpful addition to premium, letting people use their own choice of agent to help them construct complex arguments to assess.
The premium option needs a lot of work. It's weak in the import/export area. And it needs to be much more social. I have ideas for making it such that one can find contrasting opinions from any opinion page. I imagine people of diverse opinions coming together to find common ground and understand each other better. It might not be as hard as commonly thought. I'm going to try.
Closing thoughts
I watched Twitter succeed when its main distinguishing characteristic was that it limited posts to 140 characters. This totally non-novel, small design decision had surprise effects by creating a platform where everything was concise. I copied Twitter in this respect, but of course, as a child of the 8-bit computing area, I used the one true length limit of 255.
HowTruthful also has its own small design decisions that aren't particularly novel technologically, but have nice surprise effects. The default value of "debatable" as you enter new statements starts to feel really nice when you're brainstorming, for example.
But perhaps the most surprising was when I compromised on the "make everything as clean and minimialistic as possible" principle I mostly follow, and removed the code that hid the "Con" section header when it was empty. Actually it was both "Pro" and "Con", but the nice surprise effect was all about "Con".
Leaving the "Con" header floating there over empty space turned out to be a reminder to always consider contrasting opinions. So I'll close doing that.
My opinion is that there's tremendous value in the dispassionate reasoning facilitated by structuring arguments the way HowTruthful does. A contrasting opinion is that there's value in passionate prose. It's always entertaining when passionate prose emerges unexpectedly from an LLM. Of course, the LLM itself has no passion. It doesn't even have an underlying meaning to its words. Any "meaning" in its model is just relative to other opaque words. But the prose that pops out is a statistical prediction based on everyone whose thoughts went into its training data.
So when I was doing my due diligence and having Claude review the argument map I'd made about whether this submission was a winner, all those people whose thoughts went into the training data took me to task severely for underselling it. I didn't want to translate their points into a graph of dispassionate context-independent statements. I can't believe I'm saying this, but please read this AI-generated text: claude-assessment.md