Three months before the review
The assurance of learning committee is looking at the evidence file for a goal every business school in the world has some version of: graduates communicate effectively in professional settings.
What's in the folder is a presentation rubric from the capstone, scored on organization, delivery, and visual design. A peer-evaluation instrument from a team project. A reflection paper from the ethics course. An employer survey with a 14% response rate, two years old.
Nobody in the room thinks this measures whether graduates can communicate effectively in professional settings. The associate dean knows it. The faculty on the committee know it. It goes in the folder anyway, because the alternative is an empty folder, and because everyone has been doing it this way for fifteen years.
This isn't a failure of rigor. Business faculty are perfectly capable of designing a valid measure. The problem is that until very recently, the valid measure was not physically possible to administer, and everyone quietly agreed to measure something adjacent instead.
That constraint has now changed, which makes the workaround harder to justify.
What the current instruments actually measure
It's worth being precise about the failure modes, because they're different from each other and each one is instructive.
Written work measures knowledge of the behavior, not the behavior. A student who writes an excellent memo explaining why a junior auditor should escalate a suspected fraud indicator has demonstrated that they know escalation is correct. That was never in doubt. Whether that same student can say it out loud to a partner who is visibly leaving for the day is a separate capacity, and the correlation between the two is weak enough that treating one as evidence of the other is indefensible when stated plainly.
Presentation rubrics measure delivery, not judgment. Organization, pace, eye contact, slide quality. A student can score in the top band while making every substantive error available — burying the recommendation, conceding the wrong point, mistaking a hostile question for an attack. The instrument is well-designed for what it measures. It just doesn't measure the thing the learning goal names.
Peer evaluations measure reciprocity and likeability. Very few faculty will defend them as a valid measure of teamwork or leadership. They persist because a team project is the only place per-student interpersonal data appears at all, and something is felt to be better than nothing.
Self-report measures confidence, and in exactly these skills confidence tends to run inverse to competence. The student most certain they handled the difficult conversation well is frequently the one who hedged the message into ambiguity and left believing it landed. That is not a marginal measurement error; it points the wrong way.
Indirect measures — employer and alumni surveys — measure perception, at low response rates, with a multi-year lag, unattached to any course, cohort, or intervention. Useful context. Not a measure you can close a loop with.
Put together: a program can hold a well-documented, carefully rubricked assessment process for communication, leadership, and ethical reasoning in which not one instrument observes a student communicating, leading, or making an ethical decision.
The constraint was never the rubric
Here's the part that gets misdiagnosed, and it matters for what to do about it.
Faculty have always known how to assess a behavior: put the student in the situation, watch what they do, score it against criteria. That's a solved problem in nursing, in medicine, in aviation, in law school trial advocacy. The rubric was never the obstacle.
The obstacle was the counterparty.
To observe whether a student can hold a position under pressure, someone has to apply the pressure — credibly, consistently, and to every student. And that person has to be a hostile CFO, or a skeptical engagement partner, or a procurement officer trained to commoditize, or an employee disclosing that they're struggling.
Faculty can play those roles for a handful of students. In a section of forty, a fifteen-minute conversation each consumes ten hours. Across four sections it's a full teaching load for one assessment. Practitioner volunteers help, don't scale, and vary enormously from one alum to the next — which destroys the comparability the assessment depends on. Peers cannot generate the pressure at all: a classmate has no authority, no expertise in the counterpart's role, and every social reason to be gentle with someone they'll see again on Thursday.
So programs measured what could be measured at scale. That was a reasonable response to a real constraint.
The constraint is what has changed. Conversational simulation supplies a counterparty who behaves consistently, applies real pressure, is available to every student in every section, and produces a transcript. It doesn't supply the pedagogy or the rubric — faculty still write those, and they're the same ones faculty would have written all along. It removes exactly one bottleneck, and it happens to be the one that made behavioral assessment impossible.
What a direct measure looks like
The difference is most visible side by side.
Goal: graduates demonstrate ethical reasoning.
Current: a reflection paper on an ethics case, scored 1–4.
Direct: in a simulated conversation where a supervisor pressures the student to backdate a shipping record, did the student restate the request in plain terms, decline it, avoid character accusation, and identify an appropriate next step? Scored against a fixed rubric, with the transcript retained.
Goal: graduates communicate effectively under pressure.
Current: capstone presentation rubric.
Direct: in a simulated investment meeting, did the student state the recommendation in the first fifteen seconds, produce a variant perception, volunteer the disconfirming case, and recover the thread after an interruption?
Goal: graduates lead and manage others effectively.
Current: peer evaluation from a group project.
Direct: in a simulated performance conversation, did the student state the rating within the first thirty seconds, cite at least two dated specific incidents, and could the employee — at the end — state what would produce a different rating?
Goal: graduates work effectively across cultures.
Current: study-abroad participation, or a self-report instrument.
Direct: in a simulated supplier remediation conversation, did the student test a readily-given commitment for feasibility rather than accepting it, and did they acknowledge their own organization's contribution to the problem?
In each pair, the second is an observation of the behavior the goal names. The first is an observation of something correlated with it, at best.
Why this holds up in a peer review
Deans have a specific practical question: does this survive scrutiny from a visiting team? Five properties matter.
Consistency across sections and instructors. Every student faces a counterparty that behaves the same way, scored against the same rubric. This is something neither live role play nor a practitioner panel can offer, and it's usually the first thing a review team probes.
A retained artifact. The transcript is the evidence. It can be reviewed after the fact, sampled, re-scored, and shown. Most soft-skills assessment produces a number with nothing behind it.
Quantitative components where they exist. Some behaviors are directly countable — time elapsed before a termination decision was stated, the ratio of diagnosed problems to prescribed solutions in a creative review, the price a student landed against a known floor. Not everything reduces to a number, and where it does, the resulting measure is unusually defensible.
Repeatability across cycles. The same instrument can run next year against the same rubric, which is what makes a genuine trend line possible.
Course-embedded rather than bolted on. The assessment happens inside the course where the learning happens, which is what review teams prefer and what faculty will actually sustain.
Closing the loop — the part programs actually get graded on
Assurance of learning isn't really about measurement. It's about demonstrating a cycle: you measured, you found something, you changed something, you measured again, and it moved.
Soft skills are where that cycle most often stalls, because the measures are too coarse to reveal anything actionable. A reflection paper scored 3.2 out of 4 tells a curriculum committee nothing about what to change.
Behavioral data produces findings you can act on. Seventy-one percent of students buried the performance rating past minute five. That's a specific, teachable deficiency. You add a module on leading with the decision, you re-run the same simulation the following term, and the number moves — or doesn't, which is also information.
That is a closed loop with a real intervention and a real re-measurement, on a learning goal where most programs can produce neither. For a continuous improvement review, it's a substantially stronger story than any amount of process documentation.
What this doesn't do
Three honest limitations, because the argument is weaker without them and any faculty member worth having on the committee will raise them.
A simulation is still a proxy. It measures behavior in a simulated conversation, not behavior in a real one where a career is at stake. That's a better proxy than an essay about the behavior — it's not the behavior itself, and it shouldn't be described as such.
Rubric quality still governs measurement quality. A vague rubric produces vague data faster than before. The design work doesn't disappear; it becomes the whole job.
Scoring reliability moves rather than vanishes. If scoring is AI-assisted, it needs validation against faculty scoring on a sample, with disagreement examined rather than averaged away. That's a real methodological obligation and programs should plan for it rather than discover it during a review.
There's also a fourth objection worth answering directly: students will perform to the rubric. They will. But the rubric describes the behavior the program wants them to exhibit — leading with the decision, testing a commitment, declining an improper request. A student who performs those behaviors repeatedly under pressure, until they're available without deliberation, has acquired the thing the curriculum was trying to teach. Performing to the rubric is a problem when the rubric measures polish. It's the mechanism when the rubric measures judgment.
Where to start
Not a program-wide rollout. The version that works is deliberately small:
One learning goal. One course. Two sections. One term. Pick the goal with the weakest current evidence — for most programs that's ethical reasoning or interpersonal effectiveness. Build one scenario tied to it. Run it in two sections with a common rubric, retain the transcripts, and score a sample by hand alongside whatever automated scoring is used.
At the end of the term you have something almost no program has: a direct, per-student, rubric-anchored measure of a behavior, with the underlying artifact retained. That's either a better assessment or it isn't, and one term is enough to know.
The short version
Business programs are not vague about what they want graduates to be able to do. Communicate under pressure. Lead people through difficult conversations. Reason ethically when it's expensive. Work across cultures without condescension.
They have been vague about measuring it, and for a long time they had a good reason: you cannot give four hundred students a hostile CFO. Reflection papers were what remained.
That constraint is gone. Which means the question in front of a program is no longer whether behavioral assessment of soft skills is feasible. It's whether the evidence file, three months before the next review, should contain a measure of what students know about professional judgment — or a record of what they actually did when someone pushed back.
This is the concluding piece in a series examining twenty specific professional conversations — auditor and CFO, manager and employee, founder and investor, sourcing manager and supplier — that business curricula cover analytically and rarely rehearse. Foretell AI lets faculty build those conversations as simulations, with configurable counterparties, transcripts, recordings, and rubric-based evaluation. If you're strengthening assurance-of-learning evidence on communication, leadership, or ethical reasoning, we're happy to walk through how other programs have structured a first pilot.