How Agencies Steer Outcomes Without Reversing a Single Decision

In 2010, the Social Security disability allowance rate at the hearing level was 62%. By 2019 it was 45%. Same statute, same regulations, same underlying medical conditions — people kept applying with degenerative disc disease, depression, congestive heart failure, year after year. A seventeen-point drop in a decade, and no one ever issued a directive telling judges to deny more claims. No rule was published. No memo went out. Every administrative law judge could accurately say they decided each case on its merits. The aggregate moved seventeen points anyway.

This episode explains how. It's the third and final installment of a three-part arc on consistency in administrative adjudication, and it answers the question the first two set up. The first episode showed there's no horizontal mechanism pulling ALJ decisions toward each other. The second showed there's no effective vertical mechanism either — internal review corrects little and produces no consistency. Yet agency preferences clearly translate into outcomes. So something is doing the work, and it isn't reversal.

The answer is that agencies mostly don't control adjudication by reversing decisions they dislike. They control it by shaping the conditions under which decisions get made — upstream of the hearing, before the record is even developed, through case completion targets, performance evaluation, guidance documents that aren't technically law, and quality review. Reversing decisions is expensive and leaves a trail lawyers can read. Changing what the average judge thinks the right answer is costs nothing per case and leaves only an aggregate statistic with no visible cause. Nobody can appeal a training program. Nobody can sue because their judge had a quality review meeting.

The episode is careful about a real objection: most of these mechanisms are just ordinary management. Every institution trains its people and evaluates their performance, and some consistency in interpretation is exactly what the rule of law asks for. The mechanisms aren't illegitimate. The problem is cumulative — what these practices do when applied at scale to people whose job is supposed to be deciding cases independently, and the fact that the procedural protections the APA attaches to binding policy don't attach to any of them. Management practice, applied to adjudicators at scale, starts to look like policymaking without the procedure that's supposed to come with it.

Courts have seen this for forty years and struggled with what to do. In Association of Administrative Law Judges v. Heckler (1984), the Social Security ALJ union sued over a review program that targeted high-allowance judges for enhanced scrutiny; the court found it created an atmosphere of tension and unfairness that violated the spirit of the APA "if no specific provision thereof" — and being unable to name a provision made it nearly impossible to enforce. In a 2015 case over case-production quotas, Judge Posner's court declined to reach the merits, but a concurrence captured the harder problem: serious impairment of the adjudicative function can come "at the hands of officials with the most worthy of motives." You don't need bad faith. You just need the pressure.

The episode walks through the modern apparatus — case completion targets that measure throughput but not thoroughness, approval-rate tracking that pushes the opposite direction, Social Security Rulings that the agency itself describes as both lacking the force of law and binding on every component, and quality review as the transmission mechanism that converts all of it into career consequences for individual judges. A hypothetical judge, "Williams," shows how it works: a flagged favorable decision, a memo to her supervisor, a rubric she can't fully see, and a next case where the decision tilts a degree — same evidence, no directive, no reversal. Multiply that across hundreds of thousands of cases and you get the seventeen-point drop. The claimant who loses today, on a file that would have won three years earlier, can't see any of it. Neither can the lawyer.

The steering apparatus this episode describes has always operated in tension with the formal independence protections that were supposed to insulate ALJs from institutional pressure. Those protections were never sufficient, but they existed — and they existed because the Merit Systems Protection Board enforced them. In Trump v. Slaughter (June 2026), the Supreme Court overruled Humphrey's Executor and ended for-cause removal protection for most independent-agency officials. The reasoning applies to the MSPB directly. The double-shield structure that made ALJ tenure protection meaningful has functionally collapsed, and the enforcement mechanism behind ALJ tenure protection is now substantially weaker than at any time since the APA was passed in 1946. The steering mechanisms this episode catalogues are unchanged by Slaughter. What has changed is the counterweight — the formal protections that were supposed to push back against the steering are themselves under active constitutional pressure. The mechanisms this episode describes were already producing seventeen-point aggregate shifts under a system where ALJs had tenure protection. What they produce under a system where the protection has thinned is the next chapter — one that will be written case by case, hearing by hearing, over the years to come.

Listen Now

Spotify | Apple Podcasts | Listen on our site

What We Cover

  • The seventeen-point drop in Social Security disability allowance rates from 2010 to 2019 — and why no rule, memo, or directive can be found to explain it
  • Why reversing decisions is the expensive, visible way to control outcomes, and why shaping the conditions upstream of the hearing is cheaper and quieter
  • The legitimate-management objection: why training, performance evaluation, and interpretive consistency aren't the problem, and what the cumulative effect at scale is
  • Association of Administrative Law Judges v. Heckler (1984): how outcome-based targeting of high-allowance judges violated the "spirit" of the APA — and why naming no specific provision made it unenforceable
  • The 2015 production-quota case and Judge Posner's concurrence: how the adjudicative function can be impaired by officials with entirely worthy motives
  • Why case completion targets measure throughput but not record development, accuracy, or whether the claimant had a meaningful chance to be heard
  • How productivity pressure pushes toward approvals while approval-rate tracking pushes toward denials — and how judges navigate both with no one telling them which wins
  • Why Social Security Rulings function as binding policy issued without notice-and-comment, and how SSR 82-13 directs own-motion review specifically at favorable decisions
  • How quality review converts metrics and guidance into performance-evaluation consequences, illustrated step by step through a hypothetical flagged judge
  • Why immigration is the amplified version — the 2018 EOIR quota memo, weaker independence protections for DOJ-employee judges, and the Attorney General as a trump card on top
  • Why the formal independence protections for ALJs are real but protect something narrower than freedom from institutional pressure

Full Rough Transcript

Gwen: Hello, and welcome to Administrative Remedies, because you can’t fix what you don’t understand. Brought to you in part by the University of Tulsa College of Law. I’m Gwendolyn Savitz, an associate professor here at TU and the associate dean of research and intellectual life.

Marc: And I’m Marc Roark. I’m the dean of the College of Law.

Gwen: We’ll be breaking down complex doctrines with real-life analogies and examples to demystify the world of administrative law for everyone trying to understand how government actually works.

Marc: Agencies are the main way the federal government gets things done. It’s not through Congress for reasons we’ll be addressing over the course of this series.

Gwen: So, Marc, in 2010, the Social Security disability allowance rate at the hearing level was 62%. By 2019, it was 45%. Same statute, same regulations, same fundamental medical conditions people apply with. Degenerative disc disease, depression, congestive heart failure, same diagnoses year after year. But a 17-point drop in a single decade.

Gwen: And nobody issued a directive telling ALJs to deny more cases. No SSR was published saying tighten up. No memo went out. Every individual ALJ will tell you accurately that they decided each case on the merits as they understood the merits.

Marc: And the aggregate still moved 17 points.

Gwen: Right. So in the last two episodes, we talked about what doesn’t explain it. That there isn’t a horizontal mechanism. The ALJs don’t bind each other. But the operational law is privatized behind local counsel, so there’s no bottom-up pull towards consistency. And then in the last episode, there’s no effective vertical mechanism for it. The Appeals Council reviews almost nothing. The BIA is selective and political, and the SEC’s review is structurally tilted. None of those produce consistency, and none of them explains the 17-point shift.

Marc: And still the shift happened.

Gwen: Right, across thousands of ALJs, across every hearing office without any of the mechanisms we’ve been studying for the last two episodes producing it. Something else is doing the work.

Marc: Something other than just a reversal of the decisions.

Gwen: Right. Agencies mostly don’t control outcomes by reversing decisions they don’t like. They control outcomes by shaping the conditions under which the decisions get made, before the hearing, before the record is even developed, through training, metrics, guidance documents that aren’t technically law, performance evaluations, and quality review. None of those things look like correction. But all of them are how correction happens.

Marc: So they happen upstream from the hearing?

Gwen: Yes, and mostly invisible to everybody outside the agency.

Marc: Before we get into the mechanisms, though, the intuitive picture of how agencies control adjudication is exactly the one you’re saying isn’t the main channel for these changes. Agency reviews a decision, doesn’t like it, reverses it. That’s the model, I would assume.

Gwen: Right. And that kind of ex-post correction does happen. Last episode, we talked about the mechanisms that do it. They’re real. They exist. But they’re not the dominant force for a couple of reasons. The first is cost. If the agency wants to shift outcomes by reversing individual decisions, it has to review every case it wants to change. That’s expensive in a system working at any kind of scale. If it can change what the average ALJ thinks is the right answer in a particular kind of case, it doesn’t need to reverse anything. The decisions come out the way the agency wants without any individual case being touched.

Marc: Yeah, it becomes more efficient.

Gwen: And quieter. So at agencies where appellate decisions get published, like the SEC or the NLRB or the BIA’s precedential decisions, a reversal leaves a trail that lawyers can read and cite. At Social Security, even reversals don’t surface that way. The Appeals Council doesn’t issue precedential decisions. We’ve talked about that. But either way, behavioral change leaves a different kind of trail than reversal. You can see the 17-point drop, but you can’t necessarily see the mechanism behind it. Nobody can appeal a training program. Nobody can sue because their judge had a quality review meeting. The aggregate is visible, but the cause is not.

Marc: So it’s cheaper and quieter.

Gwen: Yes. The agency doesn’t necessarily have to tell people what it wants. It builds the conditions under which what it wants comes out anyway.

Marc: So let me push back on something, though. Some of what you’ve just described just sounds like ordinary management. Every organization trains its people. Every organization has performance evaluation. If I’m running anything, a law school, a company, a hospital, I want there to be consistency. I also want there to be quality, and I want my people to be aligned. That doesn’t seem sinister. That just seems like it’s running an institution.

Gwen: Yeah, that’s true. So the mechanisms themselves aren’t illegitimate. Training is necessary. It’s a good thing. Performance evaluations make sense. Some consistency in interpretation is what people mean when they’re asking for the rule of law. The question isn’t whether agencies should manage their adjudicators. They have to. The question is what the cumulative effect of these mechanisms is when they’re applied to people whose job is supposed to be deciding cases independently, and whether the formal protections that we attach to independence actually cover what’s happening.

Marc: So you’re not saying the existence of training is the problem?

Gwen: No, not the existence. It’s just trying to acknowledge the cumulative impact and this lack of transparency, what these actually do, and the fact that the procedural protections the APA built for binding agency policy don’t really attach to any of them. The mechanisms are management practices that, when applied to adjudicators at scale, do something that looks an awful lot like setting policy without any of the procedure that’s supposed to attach when agencies do it.

Marc: You said this isn’t new. How long has this been going on?

Gwen: It’s been documented for at least 40 years. There’s a case that shows courts have seen this dynamic and struggled with what to do about it. This is the Association of Administrative Law Judges v. Heckler. This is 1984. The Social Security ALJs’ union sued the Secretary of Health and Human Services. Their claim was that the agency was targeting judges with high allowance rates, the ones that were approving too many disability claims, and targeting them for enhanced quality review. Their decisions were getting flagged. Their files were under extra scrutiny. The implicit message, the union argued, was your approval rates are too high. Bring it down.

Marc: And the agency never said that out loud.

Gwen: Right. Nobody wrote a memo saying approve fewer claims. What happened was the review apparatus pointed itself at specific judges based on outcome. The court ruled against the union on most issues, but this part didn’t.

Marc: Defendants’ unremitting focus on allowance rates in the individual ALJ portion of the Bellmon Review Program created an untenable atmosphere of tension and unfairness which violated the spirit of the Administrative Procedure Act, if no specific provision thereof. Defendants’ insensitivity to that degree of decisional independence the APA affords to administrative law judges and the injudicious uses of phrases such as targeting, goals, and behavior modification could have tended to corrupt the ability of administrative law judges to exercise their independence in the vital cases that they decide.

Marc: “Could have tended to corrupt.” The court is saying that targeting based on outcomes is the kind of pressure that bends decision-making, even where you can’t trace any single decision to it.

Gwen: Right. And then the last phrase of that first sentence is talking about the structural problem, that it violated the spirit of the APA, even if it didn’t violate any specific provision. The court’s not pointing to a specific APA section, it’s saying that the APA was written for a particular kind of agency action, rulemaking, adjudication, these defined processes with defined procedural protections. What the agency was doing here, where it was applying review pressure in patterns that tracked outcomes, didn’t really fit cleanly into either category. So the court could say it violated the spirit of the statute, it just couldn’t say which provision, which made it nearly impossible to enforce.

Marc: And 40 years later?

Gwen: Different versions, same dynamic. That 1984 case was a cruder version. You could see the targeting. It was legible. They literally used the words targeting in agency memos, which the court called injudicious. Modern versions are more sophisticated. They operate through institutional systems that don’t single out individuals overtly and don’t produce this kind of evidence, the kind that makes litigation easy. The structure evolved, but the underlying mechanism didn’t.

Marc: Walk me through the modern version then. Let’s start with the thing I think I understand the best, case completion targets.

Gwen: Every major adjudication agency tracks throughput. Social Security has targets for how many dispositions an ALJ should produce in a year. And the official rationale is entirely defensible. There are millions of pending claims. There’s real human harm from the backlog. Claimants can wait two to three years for a hearing. The agency has a responsibility to move cases.

Marc: I’m definitely not arguing with that.

Gwen: And I’m not saying it’s illegitimate in principle. The problem is what productivity pressure measures and what it doesn’t. A case completion target captures how many cases you finish. And thoroughness and speed are in tension here. They have to be. A more careful hearing is going to take longer. A more fully developed record is going to take longer to build. An ALJ who reliably hits production may be running efficient hearings of high quality, or they may be cutting corners, and the metric doesn’t tell you which one.

Marc: So has there been a case on this directly?

Gwen: In 2015, the same ALJ union sued, this was the Association of Administrative Law Judges v. Colvin, over their case production targets. The agency had directed each ALJ to manage their docket in such a way that they will be able to issue 500 to 700 legally sufficient decisions each year. The Seventh Circuit, Judge Posner writing for the majority, declined to reach the merits. Posner held that the production quota was a personnel action and the union didn’t have a direct judicial remedy under the Civil Service Reform Act. But there is a concurrence that’s worth hearing for what it talks about as the harder problem.

Marc: Serious impairment of a governmental function can occur at the hands of officials with the most worthy of motives. The integrity of the judicial function, at any level of adjudication, can be undermined seriously by even the most benignly motivated administrative or executive action that alters the essential function of adjudication.

Marc: So he’s saying you don’t need bad faith to produce the problem. You just need the production pressure.

Gwen: Right. And that the production pressure can alter the essential function of adjudication, even if nobody means it to, which has become the pattern. Courts look at this and acknowledge that the concern is real, and then they can’t reach it because the structural way these pressures operate, through general administrative practice that doesn’t take the form of any particular agency action, doesn’t really fit the shape of a concrete legal claim. So it doesn’t get resolved and it just keeps happening.

Marc: So what you’re saying about approval rates is productivity is one thing. The agency tracking how often you say yes is something else.

Gwen: Right. And they track that, too. So an ALJ who approves a claim writes a favorable decision. Relatively short. They don’t need to defeat the evidence. It won’t be reviewed at all in most cases because nobody’s going to appeal a favorable decision. And the ALJ who denies something has to write a denial that explains why every favorable piece of evidence wasn’t enough. It’s long, careful, or should be long and careful, because it has to survive whatever review actually happens.

Marc: And so that productivity pressure pushes towards approvals.

Gwen: Right. And approval-rate pressure pushes the other way. ALJs are navigating both at the same time, often without anyone telling them which one is supposed to win.

Marc: Okay, let’s talk about guidance. You’ve said before that agencies issue a lot of it. I know that’s technically not law, but I also feel like I’m missing something about how it actually functions.

Gwen: Yes, we’ll be spending an entire episode on it soon. But you’re missing the gap between the formal description and the reality on the ground. So formally, guidance definitely isn’t law, doesn’t go through notice and comment, doesn’t have the force and effect of a regulation. If you ask an agency lawyer, they will tell you that guidance documents describe how the agency is just interpreting the law. It’s descriptive, not prescriptive.

Marc: But.

Gwen: But Social Security has a weird category called Social Security Rulings, or SSRs. And those tell a different story. They’re published in the Federal Register. They’re issued under the Commissioner’s authority. They interpret statutory provisions. And their own description of their status can’t really seem to figure out what they are. This is from SSA’s own published preface to its rulings.

Marc: Although Social Security Rulings do not have the force and effect of law or regulations, they are binding on all components of the Social Security Administration and are to be relied upon as precedents in adjudicating other cases.

Marc: So they don’t have the force and effect of the law or regulations, and they’re binding on all components, relied upon as precedents. That is all in the same sentence and feels wildly contradictory.

Gwen: It sure does. It sounds like SSA doesn’t seem to know what they are.

Marc: Yeah, this is basically binding guidance issued without notice and comment.

Gwen: Right. And the APA built notice and comment as a mechanism for agencies to make binding policy. Under that, the public sees the proposal. They have a chance to comment. The agency has to respond on the record. You can challenge a rule in court if it wasn’t properly adopted. None of that apparatus attaches to an SSR. They’re issued. They bind every ALJ in the country on substantive interpretive questions. And the procedural protections in the APA that the APA was built for for binding policy don’t apply.

Marc: Okay, give me a concrete example of what we’re talking about. I want to see what kind of work these actually do.

Gwen: All right, here’s one. SSR 82-13. That means it was issued in 1982. And officially, it’s just procedural. It establishes the Appeals Council’s program of review on its own motion. They can pull an ALJ’s case for review on their own initiative without anyone having appealed it. And the ruling tells you in its opening paragraph which decisions the program is aimed at.

Marc: The policies for own-motion review, as well as the delegation of authority to the Appeals Council, are stated herein for public understanding in view of the institution of an ongoing program of review of ALJ decisions that are particularly favorable to claimants.

Marc: So only looking at decisions particularly favorable to the claimants?

Gwen: Right. So this program, which can reach any ALJ decision within 60 days and reverse it without any party having asked for review, is directed by this policy to focus on approvals. That’s a one-directional review.

Marc: Denials don’t seem to get the same treatment.

Gwen: Right. So the agency has a substantive policy preference, tighter approval standards, and that preference is built into the architecture of review. And it’s done through a document issued without public input that functions as binding procedural policy across the entire adjudicative system.

Marc: And this is in the Federal Register.

Gwen: Yes, it’s published, so it’s technically public, but it’s practically invisible to someone who isn’t already inside this world, which is part of how this works. That formal channel is transparent, but the operational effect is an ALJ who issues approvals knows or should know that those decisions face enhanced review. That knowledge can shape behavior before the next case comes in.

Marc: And it could seem like this own-motion review, of course they’re going to have to look into the approvals because the denials would be appealed on their own, but that’s not automatically true. People won’t necessarily know to appeal a denial. And that means that they’re never going to figure out those were wrongly denied.

Marc: You keep mentioning quality review. I assumed it was just monitoring, but it sounds like it’s doing something more specific.

Gwen: Yes. Quality review is what converts everything else into behavioral change. The metrics and guidance and training are inputs. But quality review is the transmission, the thing that takes those inputs and produces professional consequences for individual ALJs.

Marc: And without it, the inputs then just float?

Gwen: Right. An ALJ who knows their case completion rate is tracked but faces no real consequences from the tracking has no reason to attend to it, like the effort that students put into pass-fail classes. Quality review is what makes the tracking bite, because quality review feeds into performance evaluations, and performance evaluations have career consequences.

Marc: So what does that look like at SSA?

Gwen: They have a dedicated review apparatus inside the Office of Hearings Operations where ALJ decisions are reviewed as part of ongoing oversight. Some routine sampling, some focused review on cases flagged by particular criteria. The review checks whether the decision followed the applicable SSRs, whether the analysis of the evidence was adequate, whether the procedural rules in HALLEX were followed.

Marc: Okay, so let’s make this concrete. What does this actually look like from the inside then?

Gwen: All right. Let’s think of an ALJ. We can call her Judge Williams. She’s been on the bench for seven years. She’s got a solid record. Her approval rate is around the office median. On Tuesday, she hears a case. The claimant’s 54, has degenerative disc disease, an MRI showing disc damage, a treating physician opinion supporting an inability to perform sedentary work, consistent testimony at the hearing. She finds the claimant disabled and writes a decision.

Marc: So a routine outcome.

Gwen: Right. And then three months later, a memo comes to her. Her decision in that case has been selected for quality review. The reviewing office found a concern. They think her analysis didn’t sufficiently engage with the functional capacity evidence in the file. The memo doesn’t reverse her decision. It’s not an appeal. The claimant’s award stands, but the memo goes to her supervisor and it goes in her file.

Marc: And so now Judge Williams knows things.

Gwen: Right. She knows this decision is flagged. She knows her supervisor knows. She knows that if another favorable decision of hers gets similarly flagged, there’s a pattern. She can’t fully see the rubric that the reviewing office applied. That rubric isn’t published the way regulation is. But she can infer from the memo that the way she handled functional capacity evidence is now on someone’s radar.

Marc: And so the next time she has a similar case.

Gwen: Right. So the next time she’s writing for the reviewer as much as for the claimant. She might spend more time on functional capacity, maybe a lot more time, or, and this is the part the system depends on, maybe she’s slightly more willing to find that the evidence doesn’t quite get the person to disabled. Same medical record, same testimony, but the decision tilts a degree.

Marc: And nobody told her to deny anything.

Gwen: Right. Nobody had to. This information moved through an institutional feedback loop, and she learned something about what the agency considers adequate and what it considers inadequate. That learning affects the next case and the case after that. Every favorable decision she writes is now potentially a draft going to a reviewer whose rubric she can’t fully see.

Marc: And the rubric isn’t a regulation. So she can’t read it. She can only infer from what gets flagged.

Gwen: Right. But the criteria can shift with administration changes. So you never stop being evaluated, but the clock can reset every time that leadership’s priorities change. And if it keeps happening, three or four flags over six months, her supervisor starts having conversations with her, not about specific cases, about patterns, and whether her approach is consistent with current agency expectations, whether she’s reading the SSRs the way the agency reads them. Those conversations can feed into her annual performance review, which affects her standing, her assignments, and her professional life inside the agency. None of it’s reversal of any specific case, but all of it shapes what she does in every subsequent case.

Marc: OK, you said earlier the APA protects ALJs from removal based on the outcome of any individual case. But this isn’t removal.

Gwen: Right. So that APA removal protection, that’s a little in the air right now. But on paper, at least, that protection is real. But it’s a narrower thing than freedom from institutional pressure. The ALJ can’t be fired for deciding the case wrong, but she might be placed under quality review that correlates strongly with certain outcomes. She can have a performance evaluation that reflects patterns in her decision-making. She can receive feedback that makes the agency’s expectations clear. None of it is removal for a case outcome, but all of it shapes how she’s handling the next few thousand cases of her career. Quality review is a permanent shadow apparatus that’s evaluating ALJ work against standards that aren’t fully public and that evolve over time. So the formal independence protections are real, assuming that they stay real. They just protect something narrower than what most people imagine.

Marc: So immigration has come up in basically every episode of this trilogy as the more aggressive version of whatever we’re describing. Would that be the same here?

Gwen: Oh, yes. So everything we’ve talked about, these metrics, guidance, performance review, quality review, all of this exists in immigration, plus structural features that aren’t present with Social Security. Let’s start with the amplification. In March 2018, the Executive Office for Immigration Review issued a memo from its director that was implementing new performance metrics for immigration judges. It was supposed to be effective October 1st of that year, and the memo said that to receive a satisfactory rating, an immigration judge had to complete 700 cases per year and have fewer than 15% of decisions remanded by the BIA or the circuit courts.

Marc: 700 cases a year. That seems like a lot.

Gwen: It sure does. So the pressure to move fast was built directly into the formal evaluation structure. The quotas were modified after litigation and eventually rescinded by the Biden administration. The underlying pressure didn’t totally disappear. It just took a less explicit form.

Marc: And the structural piece beyond the metrics?

Gwen: So immigration judges are Department of Justice employees. They are not ALJs under the APA. They don’t have the same statutory tenure protections. They don’t have the same removal standards or the same independence architecture. Their professional survival is more directly tied to keeping leadership satisfied. So the ex-ante control mechanisms that exist at Social Security exist in immigration with fewer institutional buffers between the judge and the pressure.

Marc: And we already saw from last episode that the attorney general can certify cases and rewrite doctrine from above.

Gwen: Right. So we now have explicit quotas on the front end, or at least an expectation that a lot of cases will move through. We have weaker independence protections throughout, and we have the attorney general as a trump card at the top. The cumulative pressure on an immigration judge is different in kind from the pressure on a Social Security ALJ. And the stakes in any individual case, whether someone is returned to a country where they may be killed, are about as high as stakes get in administrative law.

Marc: So is there a legal line between legitimate oversight and improper interference? Because we keep talking about this as a system, and systems, they don’t have villains. But somewhere in there, someone’s making choices. And some of those choices feel like exactly what the judge earlier called an atmosphere of tension and unfairness.

Gwen: So there’s no clean legal line. What courts have been willing to do is police the extremes. A direct command to decide a specific case a specific way? Everybody agrees that’s improper. Public targeting of individual judges based on approval rates? That’s what the 1984 court found, and courts have sometimes been willing to reach that.

Marc: And the structural version of this discussion, the metrics and guidance and performance systems?

Gwen: Those are absolutely tolerated. And the alternative would be courts monitoring how agencies manage their own internal operations, their training programs, their quality review criteria, their performance evaluation systems. Courts aren’t set up to do that. They don’t have the institutional capacity, and they haven’t been willing to claim the authority. So the line gets drawn at overt, individualized pressure, and the structural mechanisms that operate at scale are largely left alone.

Marc: Has anyone ever successfully challenged a quality review rubric? So a FOIA, an APA claim, anything, or is it just structurally unreachable?

Gwen: Practically unreachable, but not absolutely. FOIA litigation has surfaced pieces of internal review criteria over the years, particularly during that Bellmon era litigation in the 1980s, where discovery exposed those targeting memos that became the basis of that case. And the 2018 memo became public because journalists obtained it. So things do leak. But what there isn’t is a successful APA challenge to a quality review rubric itself. The structural problems we’ve talked about, there’s no specific APA provision that the practices violate. It applies to those challenges, too. The cases that have come closest both ended without injunctive relief. Plaintiffs win pieces. They just don’t win structural reform.

Marc: Which means the mechanisms producing the most behavioral change are the ones most insulated from challenge.

Gwen: The quality review rubric isn’t published the way a regulation is. The performance evaluation is internal. So from outside the agency, we see outputs. An aggregate approval rate, case completion targets when they surface in litigation, guidance documents if an ALJ happens to cite one in a decision. But you don’t see the mechanism. The mechanism is operating in the background.

Marc: Let’s circle back to Judge Williams’s case. The 54-year-old with the disc problem, he won, his decision stood. But you said something earlier I want to come back to. You said next time she has a similar case, the decision tilts a degree.

Gwen: Yes.

Marc: So somewhere out there is the next claimant, same age, same disc damage, same kind of treating physician opinion. And that one will come out differently.

Gwen: Right. That’s how we get this 17-point shift across a decade. It’s not one big event. It’s hundreds of thousands of cases where the decision tilted a degree because the conditions around the decision maker tilted a degree. The claimant doesn’t see any of that. The claimant sees one hearing, sees their ALJ and gets a decision. The decision cites the regulations. It walks through the evidence. It reaches a conclusion that’s formally correct in every visible respect.

Marc: But the case that would have been approved in the same hearing office by the same ALJ on the same evidence three years earlier?

Gwen: Right. It can get denied today. Not because anything about the case changed. Because the system around the judge changed. The quality review environment tightened. The feedback signals accumulated. The metrics pressure shifted. None of which the claimant can see. None of which the lawyer can see, even if they’re experienced. And none of which will show up in a federal court appeal.

Marc: And when they lose, the natural response is to ask what they did wrong.

Gwen: Right. What they should have said differently, what evidence they should have brought. Sometimes those are exactly the right questions, but sometimes they’re not. Sometimes the claimant did nothing wrong and the outcome reflects the system the ALJ is operating in rather than the merits of the claimant’s particular case.

Marc: And the lawyer can’t tell them that because the lawyer can’t see it either.

Gwen: Right. So if we look on paper, we can say that the formal independence of ALJs is real. The APA protections are meaningful. And formal independence is a narrower thing than freedom from institutional pressure. The outcomes ALJs produce reflect the system that they’re operating in, not just their individual judgment. When that system shifts, the outcomes shift across thousands of cases without any individual decision being reversed, without any directive being issued, and without visible mechanisms anyone can point to.

Marc: So the answer to how do agency preferences become hearing outcomes isn’t any single thing.

Gwen: Right. It’s all of these things together, upstream of the hearing and mostly invisible from outside and structured in ways that are hard to challenge.

Marc: So we spent three episodes on this, the hearing-level variation, the review layer, and now this. The picture just keeps getting messier.

Gwen: Much messier. So the textbook version is independent adjudicators, structured appellate review, judicial oversight at the back end. The operational reality is what we’ve spent three episodes talking about. The horizontal variation that gave the system different starting points across different offices. The narrow internal review that lets local patterns harden. The ex-ante apparatus, the metrics, the guidance, the quality review that does this work of shifting where the median ALJ lands. Judge Williams’s next claimant is in that number and her one after that, too. And neither one knows it.

Marc: OK, so what are we doing next week?

Gwen: Next week, we are doing a deep dive on Social Security disability. It might feel like we’ve been doing that for so many of these episodes, but next week we’ll be talking about substantively how the evaluation occurs. If you want to see what the structures we’ve been describing actually do to real people, that’s where you see them.

Marc: So that does it for today’s episode on Administrative Remedies. Thank you for joining us today. Please, if you enjoy this podcast and enjoy this episode, give us a like on Spotify, iTunes, or whatever platform you’re listening on. And be sure to tune in next time where we’ll continue to dive into the contours of administrative law, because remember, you can’t fix what you don’t understand.



Related Guides