Dynamics of behavioral risk-taking
A redesign of the Balloon Analogue Risk Task that makes the expected-value optimum reachable — and an exploratory result I did not expect and still cannot fully explain.
Working paper
A Dynamic-Hazard Balloon Task and an Exploratory Study of Behavioral Risk-Taking
Honors Seminar Report · Middle East Technical University, 2026. Advisor: Assoc. Prof. Dr. Gülşah Karakaya.
Read the paper (PDF)The gap everyone reports
What people say about their appetite for risk does not predict what they do when something is actually at stake. Frey and colleagues put 39 risk measures to more than 1,500 adults: the self-reports cohere into a reliable general factor, and the behavioral tasks barely load on it. The questionnaires agree with each other. They do not agree with behavior — and the behavioral tasks do not agree with each other either.
The usual reading is that behavioral tasks are unreliable. I think something narrower is going on, and it starts with how the most common task is parameterized.
Why the standard task cannot see calibration
In the classic BART the burst point is drawn once, uniformly, from the balloon's capacity N. That places the expected-value-optimal stop at N/2 — 64 pumps on a 128-pump balloon. Participants average around 40. Everyone sits on the same side of the optimum.
The consequence is structural. Within the range anyone actually plays, pumping more is always better, so pump count and earnings are nearly collinear — in the Frey data they correlate at about 0.93. There is no observable behavior that counts as going too far. The task indexes gross exposure, and a success score computed on it collapses into the same thing.
A task nobody can overshoot cannot distinguish someone who stopped early because they judged the odds from someone who stopped early because they are timid.
The redesign
Replace the single uniform draw with a sequential per-pump hazard: at pump k the balloon bursts with probability k/N. Risk accumulates as the balloon inflates. Survival is then the product of per-pump survivals, and for large N, S(k) ≈ exp(−k²/2N), so EV(k) = k·S(k) ≈ k·exp(−k²/2N).
Differentiating gives a closed form: k* = √N. On a 128-pump balloon the target moves from 64 to 11. The three profiles used in the study (N = 128, 32, 8) have discrete optima of 11, 5, and 2 pumps, presented in randomized order across a thirty-trial session so that an invariant strategy is penalized. Participants now land on both sides of the optimum, and overshoot, undershoot, and accuracy separate.
The instrument, the scoring engine, and the platform are all open source. The study ran on 70 participants: 53 online, and 17 in a synchronized laboratory tournament with a cash payout. Approved by the METU Human Subjects Ethics Committee (protocol 0176-ODTÜİAEK-2026).
What the data said
The analysis was deliberately domain-agnostic. I clustered the behavioral metrics with no self-report input, then trained a random forest to predict cluster membership from DOSPERT scores alone and ranked the subscales by permutation importance.
The subscale that tracked profitable, well-calibrated performance was not financial-gambling. It was social — willingness to challenge authority and voice an unpopular opinion. The domain-matched financial subscale ranked last.
This was found in the data, not predicted in advance. Everything below characterizes and stress-tests that pattern in the sample it was discovered in. It is exploratory, and the one confirmatory test — the pre-registered external replication — is reported in full, including the part that failed.
Three methods that fail differently
One estimator at n = 70 proves nothing, so I triangulated with methods whose failure modes do not overlap.
Canonical correlation on the three canonical Lejuez metrics (normalized pumps, explosion rate, money collected) gives a first axis of ρ = 0.56, with social loading +0.85 — far above any other subscale. Financial-gambling gets relegated to the second, orthogonal axis, which is the familiar diffuse propensity dimension. Notably, social dominance increases as descriptive variables are removed, the opposite of what an overfitting artifact does.

Regularized Bayesian estimation, modeled on Frey et al., puts social as the only subscale whose 95% highest-density interval excludes zero — both as a zero-order correlate of composite behavior (M = 0.43, HDI [0.23, 0.62]) and as a partial predictor (β = 0.45, HDI [0.22, 0.69]). Financial-gambling straddles zero throughout (M = 0.19, HDI [−0.08, 0.38]). Adding age, gender, and modality as controls barely moves it (β = 0.42, HDI [0.21, 0.68]).

The shape of those correlations matters more than their size. Social tracked earnings and calibration strongly but the raw explosion rate only weakly. High-social participants were not pumping until the balloon burst — they were stopping near the optimum and banking. That is a calibrated risk-taker, not a reckless one.
Exploratory factor analysis recovers the familiar split: a stated-risk factor carrying the DOSPERT subscales and a behavioral factor carrying the BART metrics, correlated at only φ = 0.25. Social is the exception — the weakest loader on the stated factor, and the only subscale to cross onto the behavioral one (λ = 0.36, everything else below 0.25). A leave-one-out reanalysis keeps social top-ranked in all seventy folds, so it is not one extreme participant.

The tests designed to break it
The lab and online cohorts differ in incentives, supervision, and peer presence all at once, so modality is confounded by construction. Rather than assert equivalence, I learned the canonical weights on the online cohort alone and projected the held-out lab cohort through them. The correlation survives at ρ = 0.32 — weaker, which is what an attenuated real effect looks like, and genuinely out of sample rather than fitted.

Then the real test. I froze the prediction in a pre-registration and ran it once against the Basel–Berlin Risk Study (N = 1,505). The basic association replicated at ρ = 0.115, HDI [0.062, 0.166], credibly above zero, and the ordering of social over financial-gambling held with nearly identical posterior certainty (0.957 against our 0.960).
The dominance did not replicate. In that dataset — the standard BART, 128-capacity balloons only — the signal fragments. Recreational risk becomes the primary driver and social drops to fourth of six.
Why that failure is the interesting part
The boundary condition traces straight back to task architecture. In the dynamic-hazard task, pump volume and earnings correlate at only r = 0.60, because overshooting the optimum is heavily penalized — earnings reward calibrated stopping. The portion of someone's earnings that risk magnitude cannot explain is what tracks the social domain.
In the classic architecture that separation does not exist. Participants under-pump so severely that high and low risk have to be defined against the cohort mean rather than the optimum, risk and success correlate at r ≈ 0.92, and any success score mechanically collapses into exposure. The social signal is not absent there — it has nowhere to show up.

Held with caution
I do not know why the social domain behaves this way. A correlation between tolerance for social friction and calibrated stopping does not establish shared machinery. Sociometer and social-safety accounts offer a candidate story — both social risk and holding a strategy under adverse feedback require absorbing immediate negative signal for a longer-run gain — and it is theoretically attractive, which is exactly why I distrust it. Higher self-efficacy, lower trait anxiety, or plain task engagement would produce the same behavioral patience. Disentangling them needs a construct-validity battery built for the purpose.
The sample is 70, culturally homogeneous, with a structural confound between the cohorts. Convergence across estimators buys internal reliability and cannot substitute for external validation. The obvious next step is a pre-registered replication at n ≥ 150 with a proper construct-validity battery.
What I will claim is narrower than the headline: for decades the stated–elicited gap has been read as evidence that behavioral tasks do not work. The more constructive reading is that the association was probably there all along, and that whether you can see it depends on whether your task rewards calibration or just exposure.