Wednesday, 26 August 2026

Deferring Slow, Fallible Work Without Losing Visibility

The problem

A user action triggers a second piece of work that is slow, fans out across many calls to an external system, and can partially fail. Done inline, it looks like this:

public async Task<Result> Handle(SetReadyCommand command, CancellationToken stopToken)
{

    var entity = await LoadAsync(command.EntityId, stopToken);

    entity.SetReady();

    await dbContext.SaveChangesWithResultAsync(stopToken);

    // Slow, fans out, and fails per-item. But it's in the same request, so a failure

    // here either rolls back or taints a transition that had legitimately succeeded.

    foreach (var item in entity.Items)
    {
        await refreshService.RefreshAsync(item, stopToken);
    }

    return Result.Success();
}


The first is latency. The work takes long enough that holding the HTTP request open for it is unacceptable, and it scales with the size of the thing being operated on, so there's no safe upper bound.


The second, and more important one, is coupling. The secondary work is best-effort: if it fails, the system is still in a valid state, the previous data remains usable, and the user's primary action should stand. But if it runs inside the same request, a failure in the secondary work either rolls back or taints the primary action that had otherwise succeeded. You can't express "this part is allowed to fail" while it shares a transaction and a response code with a part that isn't.


So the work has to leave the request thread. The moment it does, you inherit a third problem: the user no longer finds out what happened. A fire-and-forget background task that silently fails is worse than a slow request.

The solution

Three pieces, and the third is the one people skip.

The triggering request validates, does the part that must be transactional, persists it, and returns — 202 Accepted if the whole request is deferred work, or a normal success response if the deferred work is a follow-on from something that did complete. Before returning, it hands a work item to an in-process queue.

The triggering handler does the transactional part, hands the rest to a queue, and returns:

var request = new WorkRequest(
    command.EntityId,
    itemIds,
    new WorkActor(currentUser.Username,
         currentUser.UserId,
         currentUser.Name),
    currentDateTime.UtcNow);

if (!queue.TryWrite(request))
{
    return Result.Failure(Errors.AlreadyInProgress(command.EntityId));
}

return Result.Success();   // the endpoint maps this to 202 Accepted


A single background worker drains the queue, one item at a time, each in its own dependency-injection scope so every pass gets fresh scoped dependencies — a database context in particular, since the long-lived worker must not share one.

Each pass writes its own progress and outcome to a status record on the entity being worked on, and the read endpoint for that entity exposes it. That record is the contract with the client: it moves to "requested" when the pass starts and settles on a terminal outcome that distinguishes complete success, partial success, and total failure, with enough per-item detail to say what didn't work. The client polls the thing it was already going to read.

Why an in-process queue rather than a message broker
Durability costs infrastructure, configuration, and a poison-message story, so it should be bought only when losing the work actually matters.

The queue is a channel plus a dictionary of entities that already have work in flight:

internal sealed class WorkQueue
{
    private readonly ConcurrentDictionary<long, byte> _active = new();

    private readonly Channel<WorkRequest> _channel =
        Channel.CreateUnbounded<WorkRequest>(new UnboundedChannelOptions { SingleReader = true });

    public bool TryWrite(WorkRequest request)
    {
        // Claim the entity before queueing: this is a per-entity lock, not just a buffer.
        if (!_active.TryAdd(request.EntityId, 0))
        {
            return false;
        }

        if (_channel.Writer.TryWrite(request))
        {
            return true;
        }

        _active.TryRemove(request.EntityId, out _);
        return false;
    }

    public void Complete(long entityId) => _active.TryRemove(entityId, out _);

    public IAsyncEnumerable<WorkRequest> ReadAllAsync(CancellationToken stopToken) =>

        _channel.Reader.ReadAllAsync(stopToken);
    public bool TryRead(out WorkRequest? request) => _channel.Reader.TryRead(out request);
}


Here it doesn't. The work is regenerable: if a pass never runs, the previous state remains valid and the user can trigger it again. Nothing is lost that can't be recreated by pressing the button a second time. That's the test — if the answer had been "a failed pass leaves the system inconsistent" or "this is a write we're obliged to make," it would need a durable outbox or a broker instead.

The in-memory option is a bounded channel abstraction with a single reader, which gives a lock-free producer/consumer handoff without writing locking code.

The design decisions that mattered

Per-key exclusion, not just deduplication. A concurrent dictionary keyed by entity id is checked before the queue write. This is not about avoiding redundant work — it's that two overlapping passes over the same entity race on the same rows and the same status record, and the second silently overwrites the first's result. A rejected enqueue returns a clean false, which the caller turns into a conflict response rather than accepting a request it won't honour. A useful side effect is that the queue is bounded in practice by the number of distinct entities, so "unbounded" is safe.


Carrying the caller's identity across the thread boundary. Background work has no request context, so the ambient "current user" service isn't resolvable in the worker's scope. The identity must be captured on the request thread and travel with the work item. This is easy to miss and shows up later as background writes attributed to nobody.


Writing the status record must not be blocked by unrelated changes. Originally the status was saved through a tracked entity, which put unrelated columns into the concurrency check. If anything else changed on that entity while the pass ran, the final write failed and the record was left on "requested" permanently — a client polling forever. The fix was a set-only update touching exactly the status column. The general lesson: a progress record has different concurrency requirements from the data it describes, and shouldn't share an optimistic-concurrency token with it.


Every exit path must end the poll. Three of them exist and all three need handling. An unexpected exception during a pass writes a failed outcome and leaves the worker alive rather than killing it. Shutdown cancelling a pass mid-flight writes a failed outcome, but only if the record is still "requested," so it doesn't overwrite a result that already landed. Shutdown with items still queued drains them and fails each one. The releasing of the per-entity lock sits in a finally, so a crashed pass can't lock an entity out of all future work.

Re-check preconditions between units of work. The entity's state can change while a long pass is running. Checking before each item means the pass can stop cleanly, keep what it already did, and report the rest as not done — rather than finishing against stale assumptions.