An upfront disclaimer: I built and sell the app this post ends with, so ignore the last section if you will.
I've read a lot lately about calibration, forecasting tournaments, Brier scores etc, and almost all of it is about predicting the world. My question is smaller: does my confidence mean anything, on my own decisions? The obvious way to check would be to remember a decision, how sure I was, and compare with how it turned out. Well, after resolution, I already know how it turned out, so whatever confidence I "remember" already moved toward the outcome, and I can't even really tell by how much.
As an example, before taking the current job, which is half office, half remote, I had another option on the table, fully remote, most other things similar. And that's already one of the problems, I can't remember them all, just that "most" were similar. At the moment I'm both content and satisfied with the decision and I'd also make it again, but is it because I got used to it? Because the job market is tough right now and that influences how I feel? I do remember how important fully remote was for me back then, but everything else is blurry and I definitely can't remember how confident I was in this choice. Right now it feels like I was sure, but the only thing that backs that up is the fully remote part, only because it is very important to me, making the emotion strong around the topic.
So memory is out of the question, I don't think anyone here needs convincing of that. What I couldn't find was a way to keep a record that holds up on a real decision, one that takes months to play out, and what I ended up with is still soft on the measurement side, for which I'd very much love feedback on.
Fatebook
Fatebook is the closest thing I found: you write a question, set a resolve-by date, get an email when the date arrives and mark it yes/no/ambiguous yourself, and it stays private unless you share it. Manifold works the same way in public, with the creator resolving their own decisions. Metaculus is the odd one out, since its admins do the resolving and you can't write a question that lets you grade it yourself. All 3 take ranges and multiple choice too, not just a yes or no, but they all want a question with a date on it.
Think of it like this:
"Will the contract be renewed within 6 months?" is a Fatebook question, it has a date and it's either right or wrong;
"Take the contract or enjoy some free time?" isn't, there's no date, the consequences might be felt for the next year and my own sense of whether it went well keeps moving for months after I supposedly know the answer.
The second one is what I actually wanted to track, the first one is only a piece of it. What Fatebook is missing is everything in between: what I learned a few weeks later and whether it changed my confidence, how I felt about the call at resolution, whether I still felt that way a few months later still and so on. That isn't a complaint about those tools, I just wanted a bit more.
What I wanted from a record
The prediction should be locked, otherwise I'd be able to change it a bit, maybe a few times, then convince or fool myself that that's what I thought at the time.
The confidence should be a number, not "fairly sure", because "fairly sure" can't be plotted against anything. Yes, for a single decision it's false precision, I know, but the number only means something across many of them anyway, and without it, well, it's just a diary.
Having the history that happens along the way. On an 8-month decision most of the useful data is in the middle and that's the part with the highest chances to disappear, because once I know how it ended, I read every note knowing the ending, every note is either a sign I should've caught or something that didn't matter, in hindsight.
So I wanted a note every now and then. I don't really have a rule for how often, except whenever something happened, with the new info, whether it made things look better or worse and where my confidence is now. Even when nothing happened, if a few weeks passed.
Taking a second look, a while after resolution, because how I feel about a call in the first weeks and how I feel about it after a few months are sometimes completely different and if nothing asks me, I don't go back on my own.
And I wanted the following three kept apart: was the prediction right, how do I feel about it, would I do it again, because satisfaction doesn't always equal how right I was about it, nor if I'd do it again. A call can turn out exactly as predicted and still feel bad a few months later, or the other way around, and collapsing them into "was it a good decision" just grades the decision by its outcome.
I thought of doing them by hand for a while, but I never really got into it, for various reasons: nothing to remind me, nothing properly structured (sure, I can set up a spreadsheet or Notion page to my liking, but nothing really clicked). Might be a me problem, but without something nagging me I tend to not come back to it, however good the intentions were at the start.
The evidence
If you look for whether decision journals work you'll probably meet a number I met: about 19% better forecasting accuracy from journaling, credited to a study in Behavioral Science & Policy. I wanted to cite it, but I couldn't find it in the journal's archive or anywhere else except posts citing each other. Maybe it exists somewhere, I only mention it because it's out there.
What I did find mostly holds up, with one exception I'll get to, so here's the short-ish version.
Fischhoff and Beyth did the obvious experiment back in 1972, before Nixon's trips to China and the USSR: they asked students to put probabilities on what would come out of them (would Nixon meet Mao, would the US recognize China, that kind of thing) and weeks/months later they asked the same students to recall the probabilities they gave. The recalled ones had moved toward whatever the students believed had happened, exactly what I described above with the job, and they couldn't tell by how much either (Fischhoff and Beyth, 1975).
The other well known one is the overconfidence test: give a low and a high guess for some number (how many countries are in the UN, let's say) such that you're 90% sure the real one is between them; the real one should fall outside about 1 time in 10, but it falls outside 4 to 6 times in 10. Russo and Schoemaker got that from a couple of thousand professionals, in Sloan Management Review, and the questions from the professionals' own industry came out about as badly as the general ones; Alpert and Raiffa had asked for 98% ranges a decade earlier and got about the same 4 in 10, where it should've been 2 in 100.
Then there's the feedback, which fixes some of it: Lichtenstein and Fischhoff showed people their hit rate after each round of confidence judgments, back in 1980, and the stated confidence moved toward reality, most of the gain after the first round. Weather forecasters are the usual example, they put a probability on rain every single day, they get scored on it the next day and they end up almost perfectly calibrated (Murphy and Winkler, 1984). And training, a bit: a module on probabilistic reasoning that takes under an hour improved Brier scores by 6 to 11% over the control group, across the 4 years of Tetlock's Good Judgment Project (Chang et al., 2016).
The last one above turned out to be the exception, though. A 2025 reanalysis took the teaming and training effects from that tournament's first 2 years (Mellers et al., 2014) and ran them through a model controlling for which questions people picked, when they answered and how hard the questions were, none of which the original design controlled, and the effects shrank, went away, or in places reversed (Hauenstein et al., 2025). It only covers the first two years, while Chang's numbers run across 4, but it's the same tournament and the same design underneath, and I'd take the training result with a grain of salt.
Of course, none of these tested decision journaling as a whole, only its parts (forecasting, calibration, feedback, etc). What I'm least confident about (pun intended) is whether the forecasting results, from geopolitical tournaments with questions that have a clear yes or no and thousands of people answering the same ones, apply to questions like "should I take this job"; I'm assuming they do, for now. The calibration training only partly carries over to other kinds of questions, from what I could find, and journaling has a selection problem too, because whoever keeps one already cares about their judgment. I wrote the longer version up in a guide, with the same links, if anyone wants to check it.
The soft parts of the measurement
No matter the tool, I found I still have a few problems I don't have clear answers to, since they're in the practice itself and not in any app.
Self-grading
Resolution is self-graded, so I decide whether my own call was right and nobody else looks at it. A second person with read access would fix that, but if I know someone else will read my prediction, I'll write it knowing I'll be held accountable, probably not 100% true. Vague predictions have the same problem, "this will probably work out" never counts as a miss, and once I start caring about the hit rate, it starts to pay to write them like that, which is Goodhart's law, more or less.
Do I grade myself honestly? Do I count the vague ones as misses? Would I write the same prediction if I knew someone was going to read it?
Small numbers
Meaningful decisions don't come often, maybe a handful a year, and calibration only shows up across many predictions, so a personal record stays too small to tell me anything for a long time, at a handful a year the first 10 resolved are a couple of years away, and I still read it as if it told me something.
The thing I built
I'm an iOS developer, so, naturally I decided to build an iOS app that runs those 4 requirements (and my 5th woven in):
a decision with a prediction, a confidence percentage and a review date
check-ins in between that keep a history of what happened, tagged positive or negative, and what it moved my confidence to
a resolution with an outcome, right, wrong or mixed, and a satisfaction score
a follow-up asking whether I'd make the same call (60 days after resolution by default but fully configurable)
It has reminders for each of them, the main thing a by-hand version can't do.
It draws a reliability diagram across resolved decisions: confidence at decision time grouped in steps of 10 (70 to 79%, 80 to 89% and so on), against how often the calls in each group came out right, and keeps a "preview" label on it until 10 have resolved.
After that, it also shows a hit rate, the mean gap between confidence and outcome, and a Brier score, about which I'm not so sure: it's a proper scoring rule, but I'm the one grading the outcomes it scores, so the precision is only partly real, and in practice I look at the groups and their counts before I look at scoring.
Pretty much most of it is read-only once written: the prediction, the confidence, the check-ins and the resolution. The title can be renamed, but the original is kept, and the confidence can move, but only through check-ins, so the movement is part of the history entries.
It's called Reckon (App Store, site), $3.99 once, no subscription, iPhone and iPad, with a Mac version I'm still working on. No accounts required, it only syncs through iCloud and there is no data collection or tracking. The per-category breakdown only appears at 15 resolved with at least 3 per category, so most of the insights take a while to show up, but it's what I thought is a good default. The calibration screen looks like this:
Sample data, so it has enough entries to display
Open questions
How a mixed outcome should score. A resolution can be right, wrong or mixed, and a mixed one counts as 0.5 in the Brier score and in the confidence-to-outcome gap, but as a miss in the per-group hit rate on the diagram. Each made sense, taken individually, but I also feel it should behave the same. Should it?
Whether the follow-up should feed the calibration numbers at all. Right now it's kept separate: "would you make this call again" doesn't retroactively mark the prediction wrong, since that would grade the prediction by how I feel about the call. But if I answer no, that says something about the original call too, and keeping them fully separate might not be right either. Should a "no" count for something, and how much?
Whether 10 resolved decisions is far too low of a bar to draw a reliability diagram at. Spread over the 10-point groups, that's about 1 per group, and at that size a single call moves a group's hit rate by up to 100 points, so the curve is mostly noise at that point, it just doesn't look like it. A per-group minimum would be the alternative, or no curve at all and only show the counts, but I had to pick a number and I did it mostly by feel. Is 10 far too low, and if so, what would you pick?
I'd love to hear about any of the 3, and about anything else in the measurement that looks wrong to you, but also any other feedback. And if you've kept a record like this in the past, I'd really like to know whether you had success with it and what where its shortcomings.
An upfront disclaimer: I built and sell the app this post ends with, so ignore the last section if you will.
I've read a lot lately about calibration, forecasting tournaments, Brier scores etc, and almost all of it is about predicting the world. My question is smaller: does my confidence mean anything, on my own decisions? The obvious way to check would be to remember a decision, how sure I was, and compare with how it turned out. Well, after resolution, I already know how it turned out, so whatever confidence I "remember" already moved toward the outcome, and I can't even really tell by how much.
As an example, before taking the current job, which is half office, half remote, I had another option on the table, fully remote, most other things similar. And that's already one of the problems, I can't remember them all, just that "most" were similar. At the moment I'm both content and satisfied with the decision and I'd also make it again, but is it because I got used to it? Because the job market is tough right now and that influences how I feel? I do remember how important fully remote was for me back then, but everything else is blurry and I definitely can't remember how confident I was in this choice. Right now it feels like I was sure, but the only thing that backs that up is the fully remote part, only because it is very important to me, making the emotion strong around the topic.
So memory is out of the question, I don't think anyone here needs convincing of that. What I couldn't find was a way to keep a record that holds up on a real decision, one that takes months to play out, and what I ended up with is still soft on the measurement side, for which I'd very much love feedback on.
Fatebook
Fatebook is the closest thing I found: you write a question, set a resolve-by date, get an email when the date arrives and mark it yes/no/ambiguous yourself, and it stays private unless you share it. Manifold works the same way in public, with the creator resolving their own decisions. Metaculus is the odd one out, since its admins do the resolving and you can't write a question that lets you grade it yourself. All 3 take ranges and multiple choice too, not just a yes or no, but they all want a question with a date on it.
Think of it like this:
The second one is what I actually wanted to track, the first one is only a piece of it. What Fatebook is missing is everything in between: what I learned a few weeks later and whether it changed my confidence, how I felt about the call at resolution, whether I still felt that way a few months later still and so on. That isn't a complaint about those tools, I just wanted a bit more.
What I wanted from a record
The prediction should be locked, otherwise I'd be able to change it a bit, maybe a few times, then convince or fool myself that that's what I thought at the time.
The confidence should be a number, not "fairly sure", because "fairly sure" can't be plotted against anything. Yes, for a single decision it's false precision, I know, but the number only means something across many of them anyway, and without it, well, it's just a diary.
Having the history that happens along the way. On an 8-month decision most of the useful data is in the middle and that's the part with the highest chances to disappear, because once I know how it ended, I read every note knowing the ending, every note is either a sign I should've caught or something that didn't matter, in hindsight.
So I wanted a note every now and then. I don't really have a rule for how often, except whenever something happened, with the new info, whether it made things look better or worse and where my confidence is now. Even when nothing happened, if a few weeks passed.
Taking a second look, a while after resolution, because how I feel about a call in the first weeks and how I feel about it after a few months are sometimes completely different and if nothing asks me, I don't go back on my own.
And I wanted the following three kept apart: was the prediction right, how do I feel about it, would I do it again, because satisfaction doesn't always equal how right I was about it, nor if I'd do it again. A call can turn out exactly as predicted and still feel bad a few months later, or the other way around, and collapsing them into "was it a good decision" just grades the decision by its outcome.
I thought of doing them by hand for a while, but I never really got into it, for various reasons: nothing to remind me, nothing properly structured (sure, I can set up a spreadsheet or Notion page to my liking, but nothing really clicked). Might be a me problem, but without something nagging me I tend to not come back to it, however good the intentions were at the start.
The evidence
If you look for whether decision journals work you'll probably meet a number I met: about 19% better forecasting accuracy from journaling, credited to a study in Behavioral Science & Policy. I wanted to cite it, but I couldn't find it in the journal's archive or anywhere else except posts citing each other. Maybe it exists somewhere, I only mention it because it's out there.
What I did find mostly holds up, with one exception I'll get to, so here's the short-ish version.
Fischhoff and Beyth did the obvious experiment back in 1972, before Nixon's trips to China and the USSR: they asked students to put probabilities on what would come out of them (would Nixon meet Mao, would the US recognize China, that kind of thing) and weeks/months later they asked the same students to recall the probabilities they gave. The recalled ones had moved toward whatever the students believed had happened, exactly what I described above with the job, and they couldn't tell by how much either (Fischhoff and Beyth, 1975).
The other well known one is the overconfidence test: give a low and a high guess for some number (how many countries are in the UN, let's say) such that you're 90% sure the real one is between them; the real one should fall outside about 1 time in 10, but it falls outside 4 to 6 times in 10. Russo and Schoemaker got that from a couple of thousand professionals, in Sloan Management Review, and the questions from the professionals' own industry came out about as badly as the general ones; Alpert and Raiffa had asked for 98% ranges a decade earlier and got about the same 4 in 10, where it should've been 2 in 100.
Then there's the feedback, which fixes some of it: Lichtenstein and Fischhoff showed people their hit rate after each round of confidence judgments, back in 1980, and the stated confidence moved toward reality, most of the gain after the first round. Weather forecasters are the usual example, they put a probability on rain every single day, they get scored on it the next day and they end up almost perfectly calibrated (Murphy and Winkler, 1984). And training, a bit: a module on probabilistic reasoning that takes under an hour improved Brier scores by 6 to 11% over the control group, across the 4 years of Tetlock's Good Judgment Project (Chang et al., 2016).
The last one above turned out to be the exception, though. A 2025 reanalysis took the teaming and training effects from that tournament's first 2 years (Mellers et al., 2014) and ran them through a model controlling for which questions people picked, when they answered and how hard the questions were, none of which the original design controlled, and the effects shrank, went away, or in places reversed (Hauenstein et al., 2025). It only covers the first two years, while Chang's numbers run across 4, but it's the same tournament and the same design underneath, and I'd take the training result with a grain of salt.
Of course, none of these tested decision journaling as a whole, only its parts (forecasting, calibration, feedback, etc). What I'm least confident about (pun intended) is whether the forecasting results, from geopolitical tournaments with questions that have a clear yes or no and thousands of people answering the same ones, apply to questions like "should I take this job"; I'm assuming they do, for now. The calibration training only partly carries over to other kinds of questions, from what I could find, and journaling has a selection problem too, because whoever keeps one already cares about their judgment. I wrote the longer version up in a guide, with the same links, if anyone wants to check it.
The soft parts of the measurement
No matter the tool, I found I still have a few problems I don't have clear answers to, since they're in the practice itself and not in any app.
Self-grading
Resolution is self-graded, so I decide whether my own call was right and nobody else looks at it. A second person with read access would fix that, but if I know someone else will read my prediction, I'll write it knowing I'll be held accountable, probably not 100% true. Vague predictions have the same problem, "this will probably work out" never counts as a miss, and once I start caring about the hit rate, it starts to pay to write them like that, which is Goodhart's law, more or less.
Do I grade myself honestly? Do I count the vague ones as misses? Would I write the same prediction if I knew someone was going to read it?
Small numbers
Meaningful decisions don't come often, maybe a handful a year, and calibration only shows up across many predictions, so a personal record stays too small to tell me anything for a long time, at a handful a year the first 10 resolved are a couple of years away, and I still read it as if it told me something.
The thing I built
I'm an iOS developer, so, naturally I decided to build an iOS app that runs those 4 requirements (and my 5th woven in):
It has reminders for each of them, the main thing a by-hand version can't do.
It draws a reliability diagram across resolved decisions: confidence at decision time grouped in steps of 10 (70 to 79%, 80 to 89% and so on), against how often the calls in each group came out right, and keeps a "preview" label on it until 10 have resolved.
After that, it also shows a hit rate, the mean gap between confidence and outcome, and a Brier score, about which I'm not so sure: it's a proper scoring rule, but I'm the one grading the outcomes it scores, so the precision is only partly real, and in practice I look at the groups and their counts before I look at scoring.
Pretty much most of it is read-only once written: the prediction, the confidence, the check-ins and the resolution. The title can be renamed, but the original is kept, and the confidence can move, but only through check-ins, so the movement is part of the history entries.
It's called Reckon (App Store, site), $3.99 once, no subscription, iPhone and iPad, with a Mac version I'm still working on. No accounts required, it only syncs through iCloud and there is no data collection or tracking. The per-category breakdown only appears at 15 resolved with at least 3 per category, so most of the insights take a while to show up, but it's what I thought is a good default. The calibration screen looks like this:
Sample data, so it has enough entries to display
Open questions
I'd love to hear about any of the 3, and about anything else in the measurement that looks wrong to you, but also any other feedback. And if you've kept a record like this in the past, I'd really like to know whether you had success with it and what where its shortcomings.