Can an Apartment Search Agent Call the Model Fewer Times and Still Find Good Matches?


In their paper FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, Lingjiao Chen, Matei Zaharia, and James Zou describe three ways to cut the cost of calling a large language model (LLM). They call these prompt adaptation, LLM approximation, and LLM cascade. The authors report that their cascade tries cheaper models first. On the tasks they tested, it can match the performance of the best individual model with up to 98 percent cost reduction. That number belongs to their tasks, so an agent that reads rental listings needs its own measurement.

I built a small agent I could inspect call by call. It checks apartment listings against a renter’s requirements, such as a two bedroom in Austin under $1,500 that allows dogs. My first version sent every listing to a strong model, even when a city name or a rent figure already ruled the listing out. When I opened the trace, that meant 2,500 model calls to find 101 real matches.

That first version is a straw man, so this article measures three starting points instead of one. The first sends every pair to the model. The second is an agent that already searches by city. The third runs a database query on city, bedrooms, and rent before any model sees a listing, which is what a careful engineer would build first.

This article shows what I ran. I took a fixed sample of 500 listings from a public dataset of 2019 United States rental ads. I wrote five renter requirements and scored every version of the agent against the same answers. The strong model was OpenAI’s gpt-6-sol, which I call Sol, and the cheaper model was gpt-6-luna, which I call Luna.

I traced every call with Weights & Biases (W&B) Weave, the W&B tool that records each step of an application. The complete script is in the article, and the logged tables are in a W&B Report.

Against the database query starting point, the final version cost about 25 times less. The estimated cost was $0.008 against $0.20 for the same 2,500 checks, and the final version returned all 101 matches with no wrong ones. Letting code settle the pets field made up about 44 percent of that drop, and using the cheaper model with a strong backup made up about 39 percent. A shorter prompt and reused answers made up the rest.

The totals were not the most useful part of the test. Three findings surprised me, and I kept all of them.

  • Weave’s cost column showed $0.0000 for every call. The real token counts were stored in the raw record, but the usage table displayed zero.

  • Sol missed a match 10 times out of 10 with the full listing, and found it 10 times out of 10 with only the title and body. Sending less text gave a better answer.

  • Both models reported a confidence of 5 on almost every answer. A rule that escalates below 5 almost never fires, so it cannot do much work.

The article also says where the test is weak. Only one of the 101 matches depends on reading text, so the test measures cost far better than it measures understanding.

What does the agent read, and how did I build the test?

The agent answers one yes or no question for each listing and each renter requirement. I call one listing checked against one requirement a pair, so 500 listings and 5 requirements make 2,500 pairs.

The listings come from the Apartment for Rent Classified dataset in the University of California, Irvine (UCI) Machine Learning Repository. It holds 10,000 United States rental ads with 22 fields, including the title, the body text, amenities, bedrooms, bathrooms, rent, square feet, city, state, and pets allowed. The dataset is licensed under Creative Commons Attribution 4.0, and its listing timestamps run from September to December 2019. These are historical ads, so nothing here describes today’s rents.

The sample has 500 listings chosen with a fixed random seed of 42. It holds 200 from Austin, 100 from Dallas, 100 from Houston, and 100 from other cities. Five renter requirements run against it, and every requirement also asks for a place that allows dogs.

  • Austin, 1 bedroom, rent up to $1,300.

  • Austin, 2 bedrooms, rent up to $1,800.

  • Austin, 2 bedrooms, rent up to $1,500.

  • Dallas, 2 bedrooms, rent up to $1,600.

  • Houston, 1 bedroom, rent up to $1,200.

The mix of cities is deliberate, because a feed that covers several cities gives a filter something to reject. It also flatters the filter, and I return to that limit in the results. The prompts and the traces leave out the street address, the coordinates, and the source listing ID. Only a row number such as row_3445 identifies a listing.

How do I know a match is correct, and what counts as cost?

A pair is a true match when the listing is in the right city, has the right number of bedrooms, costs no more than the rent limit, and allows dogs. The first three checks use fields with explicit values. The pets field says things like “Cats,Dogs” or “Cats”, and that decides the fourth check whenever the field is filled in.

The pets field is blank on 4,163 of the 10,000 listings, so the text has to decide those. There are 128 sampled listings that pass the city, bedroom, and rent checks with a blank pets field. An AI assistant (Claude) read every title and body of those listings against one written rule. The rule says dogs are allowed only when the text says dogs or pets are allowed. Silence, or a hint that does not say so directly, counts as not confirmed.

The result was lopsided. One listing, row_3445, calls itself “a pet friendly community” and adds a 25 pound weight limit, so I count it as a match. Two listings hint at pets without saying so. row_1808 lists a “Pet park” and row_8465 mentions “a pet bar for your favorite 4 legged friends”, so both count as not confirmed. The other 125 listings say nothing about pets.

Before this reading, I had written a keyword search for phrases such as “pet friendly” and “no pets”. It agreed with the reading on all 128 listings.

Two limits follow from this. The labels are one reader’s judgment, and the two hint listings are arguable. Also, only 1 of the 101 true matches depends on reading text, because the other 100 come from the pets field. The test therefore says much more about cost than about how well a model understands rental ads.

Precision and recall score the result. Precision is the share of returned matches that are correct, and higher is better. Recall is the share of true matches that were returned, and higher is better. With 101 true matches, one miss moves recall by about one percentage point.

Cost here means API inference cost, which is what the provider charges for tokens. A token is a small chunk of text, roughly a short word, and providers bill input tokens and output tokens at different rates. Output tokens include the hidden reasoning tokens a model spends thinking before it answers. The OpenAI pricing page listed these standard rates on September 27, 2026. Sol cost $2.00 per million input tokens and $10.00 per million output tokens, and Luna cost $0.10 and $0.50.

Every dollar figure below is an estimate from those rates and the token counts each response reported. It is not an invoice, and it leaves out infrastructure and labor. I also report model time, which is the sum of the time every call took.

How do I run the experiment?

The steps below work on macOS and Linux. On Windows, activate the environment with .venvScriptsactivate. I ran everything with Python 3.9.6 and these package versions, weave 0.52.17, openai 2.48.0, pandas 2.3.3, and py7zr 1.0.0. The script needs an OpenAI API key with billing enabled and a free W&B account with its API key. Without the W&B key, the script stops when it calls weave.init, so the traces and the report are unavailable.

Read Also:  Rethinking Data Science Interviews in the Age of AI

Create the working folder, install the packages, and set both keys in the terminal you will use for every command in this article.

mkdir apartment-agent-cost && cd apartment-agent-costpython3 -m venv .venvsource .venv/bin/activatepip install weave openai pandas py7zrexport OPENAI_API_KEY="paste your OpenAI key here"export WANDB_API_KEY="paste your W&B key here"

The dataset downloads as a zip file that holds a compressed 7z file. These commands fetch and unpack it into a data folder, which leaves data/apartments_for_rent_classified_10K.csv in place.

mkdir datacurl -L -o data/apartments.zip "https://archive.ics.uci.edu/static/public/555/apartment+for+rent+classified.zip"python -c "import zipfile; zipfile.ZipFile('data/apartments.zip').extractall('data')"python -c "import py7zr; py7zr.SevenZipFile('data/apartments_for_rent_classified_10K.7z').extractall('data')"

Save the following script as apartment_agent.py in the apartment-agent-cost folder. It holds everything the experiment needs, including the requirements, the field rules, the ground truth labels, the prompts, the cache, and the six configurations. Comments in the code mark the parts the later sections explain.

"""Apartment search agent cost experiment (apartment_agent.py).Run from the folder that contains data/ and this file:    python apartment_agent.py baseline    python apartment_agent.py city_baseline    python apartment_agent.py field_query_baseline    python apartment_agent.py filter    python apartment_agent.py reduced    python apartment_agent.py cached    python apartment_agent.py luna    python apartment_agent.py final --threshold 5    python apartment_agent.py repeatRequires OPENAI_API_KEY and WANDB_API_KEY in the environment."""import argparseimport concurrent.futures as cfimport contextvarsimport hashlibimport jsonimport timefrom pathlib import Pathimport pandas as pdimport weavefrom openai import OpenAIDATA = "data/apartments_for_rent_classified_10K.csv"PROJECT = "apartment-agent-cost"SOL, LUNA = "gpt-6-sol", "gpt-6-luna"# USD per one million tokens (input, output). Standard short context rates, checked 2026-09-27.RATES = {SOL: (2.00, 10.00), LUNA: (0.10, 0.50)}PROMPT_VERSION = "v1"MAX_OUTPUT_TOKENS = 400MAX_ATTEMPTS = 2  # one call plus at most one retry when the reply is not valid JSONREQUIREMENTS = [    {"id": "austin_1br", "city": "Austin", "state": "TX", "bedrooms": 1, "max_rent": 1300},    {"id": "austin_2br_a", "city": "Austin", "state": "TX", "bedrooms": 2, "max_rent": 1800},    {"id": "austin_2br_b", "city": "Austin", "state": "TX", "bedrooms": 2, "max_rent": 1500},    {"id": "dallas_2br", "city": "Dallas", "state": "TX", "bedrooms": 2, "max_rent": 1600},    {"id": "houston_1br", "city": "Houston", "state": "TX", "bedrooms": 1, "max_rent": 1200},]# ---------------------------------------------------------------- data and labelsdef load_sample(seed=42):    """Fixed evaluation sample: 200 Austin, 100 Dallas, 100 Houston, 100 other listings."""    df = pd.read_csv(DATA, sep=";", encoding="cp1252")    df["row_key"] = ["row_%d" % i for i in df.index]    parts = []    for city, n in [("Austin", 200), ("Dallas", 100), ("Houston", 100)]:        parts.append(df[df.cityname == city].sample(n, random_state=seed))    rest = df[~df.cityname.isin(["Austin", "Dallas", "Houston"])]    parts.append(rest.sample(100, random_state=seed))    sample = pd.concat(parts).sort_index()    sample["split"] = ["dev" if int(k.split("_")[1]) % 2 == 0 else "test" for k in sample.row_key]    return sample.set_index("row_key", drop=False)# Labels for listings whose pets field is blank and that pass the hard field checks for at least one# requirement. Rubric: true only when the text says dogs or pets are allowed. Silence, or a hint such as# "Pet park" (row_1808) or "a pet bar for your favorite 4 legged friends" (row_8465), counts as not confirmed.# The other 125 listings in this group say nothing about pets.BLANK_PETS_DOGS_ALLOWED = {"row_3445"}  # "a pet friendly community", weight limit 25 poundsdef dogs_ok(row):    """Ground truth pet rule. The pets field wins. When it is blank, the reviewed labels above decide."""    if isinstance(row.pets_allowed, str):        return "Dogs" in row.pets_allowed    return row.row_key in BLANK_PETS_DOGS_ALLOWEDdef hard_fields_ok(row, req):    return (row.cityname == req["city"] and row.state == req["state"]            and row.bedrooms == req["bedrooms"] and row.price <= req["max_rent"])def label(row, req):    return bool(hard_fields_ok(row, req) and dogs_ok(row))def prefilter(row, req):    """Deterministic checks. Returns reject, accept, or ask (the model must read the text)."""    if not hard_fields_ok(row, req):        return "reject", "hard_fields"    if isinstance(row.pets_allowed, str):        return ("accept" if "Dogs" in row.pets_allowed else "reject"), "pets_field"    return "ask", "pets_text"# ---------------------------------------------------------------- prompts and model callsdef requirement_text(req):    return ("Apartment in %s, %s. %d bedroom(s). Monthly rent at most $%d. Dogs allowed."            % (req["city"], req["state"], req["bedrooms"], req["max_rent"]))def full_listing(row):    pets = row.pets_allowed if isinstance(row.pets_allowed, str) else "not listed"    return ("Title: %snBody: %snAmenities: %snBedrooms: %snBathrooms: %snRent: $%sn"            "Square feet: %snCity: %s, %snPets field: %s"            % (row.title, row.body, row.amenities if isinstance(row.amenities, str) else "not listed",               row.bedrooms, row.bathrooms, row.price, row.square_feet, row.cityname, row.state, pets))FULL_SYSTEM = ('You screen apartment listings for a renter. Reply with JSON only, like '               '{"match": true, "confidence": 5}. match is true only if the listing meets every requirement. '               'confidence is 1 (guessing) to 5 (certain). If the listing does not say whether dogs are '               'allowed, dogs are not confirmed and match is false.')PETS_SYSTEM = ('Read an apartment listing. Reply with JSON only, like {"dogs_allowed": true, "confidence": 5}. '               'dogs_allowed is true only if the text says dogs or pets are allowed. Use false if it says no '               'pets, no dogs, or does not say. confidence is 1 (guessing) to 5 (certain).')client = OpenAI()LISTINGS = {}def redact(inputs):    """Weave hook. Replaces prompt text so raw listing copy never reaches uploaded traces."""    out = dict(inputs)    if "messages" in out:        out["messages"] = [{"role": m["role"], "content": "[redacted %d chars]" % len(m["content"])}                           for m in out["messages"]]    return out@weave.opdef call_model(model, kind, row_key, requirement_id):    """One traced model call. Inputs are keys only. Listing text is looked up from memory."""    row = LISTINGS[row_key]    if kind == "full":        system = FULL_SYSTEM        user = "Requirements: %snnListing:n%s" % (            requirement_text(next(r for r in REQUIREMENTS if r["id"] == requirement_id)), full_listing(row))    else:        system = PETS_SYSTEM        user = "Title: %snBody: %snPets field: %s" % (row.title, row.body,                                                        row.pets_allowed if isinstance(row.pets_allowed, str) else "not listed")    record = {"model": model, "kind": kind, "row_key": row_key, "requirement_id": requirement_id,              "prompt_tokens": 0, "completion_tokens": 0, "reasoning_tokens": 0, "latency_s": 0.0,              "attempts": 0, "valid": False, "answer": None, "confidence": None}    for _ in range(MAX_ATTEMPTS):        start = time.perf_counter()        resp = client.chat.completions.create(            model=model, response_format={"type": "json_object"}, max_completion_tokens=MAX_OUTPUT_TOKENS,            messages=[{"role": "system", "content": system}, {"role": "user", "content": user}])        record["latency_s"] += time.perf_counter() - start        record["attempts"] += 1        record["prompt_tokens"] += resp.usage.prompt_tokens        record["completion_tokens"] += resp.usage.completion_tokens        details = resp.usage.completion_tokens_details        record["reasoning_tokens"] += (details.reasoning_tokens or 0) if details else 0        try:            data = json.loads(resp.choices[0].message.content)            record["answer"] = bool(data["match"] if kind == "full" else data["dogs_allowed"])            record["confidence"] = int(data.get("confidence", 0))            record["valid"] = True            break        except (ValueError, KeyError, TypeError):            continue    return recorddef cost_usd(rec):    rate_in, rate_out = RATES[rec["model"]]    return (rec["prompt_tokens"] * rate_in + rec["completion_tokens"] * rate_out) / 1e6# ---------------------------------------------------------------- configurationsclass Cache:    """Stores pet decisions. The key includes listing text, prompt version, and model, so an edited    listing, a new prompt, or a different model never reuses an old answer."""    def __init__(self):        self.store, self.hits = {}, 0    def key(self, row, model):        digest = hashlib.sha1(("%s|%s|%s|%s" % (row.title, row.body, row.pets_allowed, PROMPT_VERSION)).encode())        return "%s:%s" % (model, digest.hexdigest())def decide(cfg, row, req, cache, threshold):    """Returns (predicted_match, list_of_call_records)."""    calls = []    if cfg == "baseline":        rec = call_model(SOL, "full", row.row_key, req["id"])        return bool(rec["answer"]), [rec]    if cfg == "field_query_baseline":        # The strongest ordinary starting point. A database query on city, bedrooms, and rent runs first.        if not hard_fields_ok(row, req):            return False, calls        rec = call_model(SOL, "full", row.row_key, req["id"])        return bool(rec["answer"]), [rec]    if cfg == "city_baseline":        # A fairer starting point. The feed is already searched by city, so only same city pairs reach Sol.        if not (row.cityname == req["city"] and row.state == req["state"]):            return False, calls        rec = call_model(SOL, "full", row.row_key, req["id"])        return bool(rec["answer"]), [rec]    verdict, _ = prefilter(row, req)    if verdict != "ask":        return verdict == "accept", calls    kind = "full" if cfg == "filter" else "pets"    model = LUNA if cfg in ("luna", "final") else SOL    key = cache.key(row, model) if cfg in ("cached", "final", "luna") else None    if key and key in cache.store:        cache.hits += 1        answer = cache.store[key]    else:        rec = call_model(model, kind, row.row_key, req["id"])        calls.append(rec)        answer = bool(rec["answer"])        if cfg == "final" and (not rec["valid"] or rec["confidence"] < threshold):            rec2 = call_model(SOL, "pets", row.row_key, req["id"])            calls.append(rec2)            answer = bool(rec2["answer"])        if key:            cache.store[key] = answer    return answer, calls@weave.opdef run_config(cfg, threshold=5, workers=8):    sample = load_sample()    LISTINGS.update({k: r for k, r in sample.iterrows()})    pairs = [(row, req) for _, row in sample.iterrows() for req in REQUIREMENTS]    cache = Cache()    started = time.perf_counter()    def work(pair):        row, req = pair        pred, calls = decide(cfg, row, req, cache, threshold)        return {"row_key": row.row_key, "requirement_id": req["id"], "split": row.split,                "label": label(row, req), "pred": pred, "calls": calls}    # Cached configurations run one pair at a time, so a repeated pet question is a guaranteed cache hit.    n_workers = 1 if cfg in ("cached", "final", "luna") else workers    with cf.ThreadPoolExecutor(n_workers) as pool:        futures = [pool.submit(contextvars.copy_context().run, work, p) for p in pairs]        results = [f.result() for f in futures]    wall = time.perf_counter() - started    out = {"config": cfg, "threshold": threshold, "n_pairs": len(results), "wall_s": wall,           "cache_hits": cache.hits, "results": results}    Path("outputs").mkdir(exist_ok=True)    name = "%s_t%s" % (cfg, threshold) if cfg == "final" else cfg    Path("outputs/%s.json" % name).write_text(json.dumps(out))    return summarize(out)@weave.opdef repeat_hard_cases(n=10):    """Asks the same question n times for the two hardest listings, to see how much answers vary."""    sample = load_sample()    LISTINGS.update({k: r for k, r in sample.iterrows()})    rows = {"row_3445": "dallas_2br", "row_8465": "dallas_2br"}    out = []    for model, kind in [(SOL, "full"), (SOL, "pets"), (LUNA, "pets")]:        for key, req_id in rows.items():            for _ in range(n):                rec = call_model(model, kind, key, req_id)                out.append(rec)    Path("outputs/repeat.json").write_text(json.dumps(out))    return len(out)def summarize(out):    rows = out["results"]    calls = [c for r in rows for c in r["calls"]]    tp = sum(r["label"] and r["pred"] for r in rows)    fp = sum((not r["label"]) and r["pred"] for r in rows)    fn = sum(r["label"] and (not r["pred"]) for r in rows)    return {"config": out["config"], "pairs": len(rows), "model_calls": len(calls),            "cache_hits": out["cache_hits"], "tp": tp, "fp": fp, "fn": fn,            "precision": tp / (tp + fp) if tp + fp else None, "recall": tp / (tp + fn) if tp + fn else None,            "cost_usd": sum(cost_usd(c) for c in calls), "model_time_s": sum(c["latency_s"] for c in calls),            "wall_s": out["wall_s"],            "mean_call_latency_s": sum(c["latency_s"] for c in calls) / len(calls) if calls else 0.0}if __name__ == "__main__":    parser = argparse.ArgumentParser()    parser.add_argument("config", choices=["baseline", "city_baseline", "field_query_baseline", "filter", "reduced", "cached", "luna", "final", "repeat"])    parser.add_argument("--threshold", type=int, default=5)    args = parser.parse_args()    weave.init(PROJECT, global_postprocess_inputs=redact)    if args.config == "repeat":        print(repeat_hard_cases())    else:        print(json.dumps(run_config(args.config, args.threshold), indent=2))

Run the configurations from the apartment-agent-cost folder, one command each. The baseline run makes 2,500 Sol calls and takes several minutes, and the others are quicker.

python apartment_agent.py baselinepython apartment_agent.py city_baselinepython apartment_agent.py field_query_baselinepython apartment_agent.py filterpython apartment_agent.py reducedpython apartment_agent.py cachedpython apartment_agent.py lunapython apartment_agent.py final --threshold 5python apartment_agent.py repeat

Each command saves its full results to an outputs folder and prints a summary. The block below is captured output from a rerun of the luna configuration in a fresh folder with the exact script above.

{  "config": "luna",  "pairs": 2500,  "model_calls": 128,  "cache_hits": 29,  "tp": 101,  "fp": 0,  "fn": 0,  "precision": 1.0,  "recall": 1.0,  "cost_usd": 0.004947700000000001,  "model_time_s": 145.34897142399993,  "wall_s": 145.813283333,  "mean_call_latency_s": 1.1355388392499994}

The models do not accept a temperature setting, so every run uses the default and answers can differ from run to run. Your numbers will be close to the ones in this article and will not match digit for digit. The rerun above already shows a difference, because it made no wrong match, and the run I captured for the comparison made one. The section on the cheaper model explains why.

Read Also:  Introducing DiffusionGemma

What does the baseline trace show?

The trace of the send every pair baseline shows one model call per pair and no sign of which calls were needed. Weave records this by wrapping a function with @weave.op, and the OpenAI integration adds a child call for every request to the model. Each trace is a tree of calls with inputs, outputs, timing, and token counts, and it works like a receipt that lists every step the agent took.

Screenshot by author, with the account name hidden. The left view shows a gpt-6-luna call whose Usage table reads 0 tokens and $0.0000, even though the stored record holds 141 input and 63 output tokens. The right view shows the prompt replaced by a redaction note, so cost figures in this article come from stored token counts and the rate card.

The left view of the screenshot contains the most surprising result of the setup. Weave counted the request correctly, yet its Usage table showed zero tokens and a total cost of $0.0000 for both models. The Weave cost documentation says Weave applies built in pricing for supported integrations. It also describes an add_cost() method for models without a price, and these runs did not use it.

A zero is not a price, so I calculated cost from the token counts stored in each call record. The one call in the screenshot used 141 input and 63 output tokens.

The right view shows the redaction. The script passes a global_postprocess_inputs function to weave.init, and that function replaces every prompt with a note such as [redacted 245 chars] before the trace uploads. I then searched every stored call in the project for distinctive listing phrases and for 400 street addresses from the sample, and found none.

That first baseline is easy to summarize. Sol made 2,500 calls, one per pair, using 618,635 input tokens and 63,782 output tokens, of which 18,424 were reasoning tokens. The estimated cost was $1.88, and the calls added up to 2,991 seconds of model time. The baseline ran eight calls at once, so its wall clock time was 376 seconds.

Is sending every listing to a model a fair starting point?

No, and a fairer comparison shrinks the savings. Sending every pair to Sol is the version I built first, and few real agents do it. Most narrow the feed before a model sees anything, so I measured two more starting points on the same 2,500 pairs with the same full listing prompt.

The second starting point sends only pairs whose city matches. It made 796 Sol calls and cost an estimated $0.61, so it is about 3 times cheaper than sending everything. The third runs a database query on city, bedrooms, and rent, then sends each surviving pair to Sol. It made 264 calls and cost an estimated $0.20, about 9 times cheaper than sending everything. That query is ordinary database work and has no model cost.

All three runs missed the same match, row_3445, which the next sections explain. I measure every later saving against the $0.20 query starting point, because a careful engineer would build it first. It still spends model calls on 264 pairs, and those pairs are where the rest of the article starts.

Which model calls can the pets field settle without a model?

Another 107 of those 264. A mail room clerk sorts envelopes by postal code, and the same clerk can also see that some envelopes already carry an approval stamp. The pets field works like that stamp. When it is filled in, it says dogs are allowed or it says they are not, and no model needs to read the listing.

A workflow diagram showing 2,500 pairs split by Python field checks into rejected, accepted, and a small group sent to a language model.
Image by author. The path of the 2,500 pairs through the final configuration, with counts from the captured run. Python settles 2,343 pairs, and the model reads only the 157 pairs whose pets field is blank.

The diagram shows where the pairs went. Field checks rejected 2,236 pairs, and 1,704 of those had the wrong city. That first branch is the same work the database query does, and it covers 89 percent of all pairs in my mixed city sample.

The pets field settled the next 107 pairs. Seven were rejected because the field said cats only, and 100 were accepted because it listed dogs. That left 157 pairs, which cover 128 different listings, for a model to read.

The filter configuration sends only those 157 pairs to Sol with the same full prompt. Its estimated cost fell from $0.20 for the query starting point to $0.11, about 43 percent lower, and its model time fell from 352 to 203 seconds. Precision stayed at 1.0 and recall stayed at 0.990, since it missed the same match as the starting points. That is what a correct rule should do, because it changes which pairs reach the model and leaves the answer for the remaining pairs alone.

Does sending less text change the answer?

It changed the answer in this test. The reduced configuration asks Sol one question, whether the title and body say dogs are allowed. It drops the amenities, rent, city, and other fields, because the field checks already handled them.

Input tokens fell from 34,194 to 23,624, about 31 percent lower, and the estimated cost fell from $0.114 to $0.098. The saving is smaller than the token drop suggests. Output tokens rose from 4,589 to 5,039, and each output token costs five times as much as an input token.

The reduced run also found the match that all three starting points missed. With the full listing, Sol answered no for row_3445, the pet friendly community, with the highest confidence. To check that this was not luck, I asked each version of the question 10 times. The full listing prompt gave 0 correct answers out of 10, and the title and body prompt gave 10 out of 10.

I do not know why. One possible reason is that the full prompt shows a line saying the pets field is not listed, right next to the instruction that unstated dogs count as not confirmed. I did not test that idea.

The repeat check also covered the second hard listing, row_8465, the one with the “pet bar”. The table below shows how many of the 10 answers were correct for each model and prompt. For row_8465, a correct answer is not confirmed, following my label.

Model

Prompt

row_3445 (dogs allowed), correct out of 10

row_8465 (“pet bar”), correct out of 10

Sol

Full listing

0

10

Sol

Title and body only

10

7

Luna

Title and body only

10

8

Both models split on row_8465, answering yes two or three times in ten. That listing is a coin flip, and the table shows how much a single run can depend on one such listing.

Can the agent reuse an answer safely?

Yes, if each stored answer remembers what produced it. A cache is a notebook of past answers, and its risk is reading an old answer for a question that has changed. The script keys each answer by the model, the listing’s title, body, and pets field, and a prompt version string. Editing a listing, changing the prompt version, or switching models therefore changes the key, and the old answer is never found.

Read Also:  Benchmarking Tabular Reinforcement Learning Algorithms

The renter’s requirement is not part of the key, because the pet question does not depend on the requirement. The question is the same when two renters both want a dog friendly place in Austin. That is why the cache helped, since the same listing shows up under more than one requirement. It answered 29 of the 157 pairs from memory, so model calls fell from 157 to 128 and the estimated cost fell from $0.098 to $0.082.

The cache in the script lives in memory for one run. A production cache needs an expiry time and a way to drop entries when a listing is removed or edited. Those parts were outside this test.

Is the cheaper model safe to use?

For this workload, almost. Luna reads the same 128 listings with the same title and body prompt, and the estimated cost fell from $0.082 to $0.005, about 16 times lower again. In the captured run Luna made one wrong match, and in a rerun it made none. It said yes to the “pet bar” listing with a confidence of 4 on a scale from 1 to 5, so precision was 0.990 and recall was 1.0.

Escalation gives the agent a second opinion. Like a junior clerk who passes uncertain cases to a senior one, Luna answers first and Sol reads again only when Luna’s confidence falls below a threshold. Luna answered 5 for 127 of its 128 questions, and Sol answered 5 for all 157 of its questions. The confidence scores barely vary, so a threshold has little to work with.

I have to be honest about the threshold. I split the listings into a development half and a test half, so I could tune the threshold on one half and report on the other. The development half had no errors from any configuration, and both hard listings fell in the test half. That left nothing to tune on, and I chose a threshold of 5 after seeing which answer had a confidence of 4. The final result is therefore an illustration and not a validated setting.

With that caveat, the final configuration escalated one answer. Sol read the “pet bar” listing again, said no with a confidence of 4, and the run had no wrong matches. It made 129 calls, 128 to Luna and 1 to Sol, and the estimated cost was $0.0079.

The rerun in the setup section adds a warning. Luna alone made no wrong match there, and it answered the same listing correctly with a confidence of 5. So one wrong match against none is noise on a coin flip listing, and my run cannot show that escalation fixed anything.

One more cost detail matters. Luna’s 128 calls used 6,105 output tokens, more than Sol’s 4,319 for the same questions, because Luna spent more tokens on reasoning. About 61 percent of Luna’s estimated cost was output tokens, so limiting reasoning effort is a lever I did not test.

A keyword search would have matched my labels on all 128 blank listings, so is a model worth keeping in this loop at all? The repeat check above shows where the two would differ, and the last section says what I would monitor to decide.

What did the combined version cost, and what does the evidence support?

The lowest cost version that made no errors in the captured runs was Luna with a Sol backup, at an estimated $0.0079 for 2,500 pairs. The query starting point cost $0.20 for the same pairs, about 25 times more, and the table lists every configuration.

Configuration

Model calls

Estimated cost

Model time

Result

Send every pair to Sol

2,500

$1.875

2,991 s

1 match missed

Same city pairs only

796

$0.613

1,072 s

1 match missed

City, bedrooms, and rent query

264

$0.199

352 s

1 match missed

Plus pets field check

157

$0.114

203 s

1 match missed

Title and body only

157

$0.098

199 s

No errors

Reuse answers

128

$0.082

168 s

No errors

Luna only

128

$0.005

131 s

1 wrong match

Luna with Sol backup

129

$0.008

139 s

No errors

The interactive version of this table, the cost chart, and every model call are in the W&B Report.

Two bar charts comparing estimated API cost and total model time for eight configurations, both on a log scale.
Image by author. Estimated API cost in US dollars, and total model time in seconds, for eight configurations on the same 2,500 pairs. The data comes from the captured runs on September 27, 2026, and both axes use a log scale. Lower is better for both charts, and the note above each cost bar says whether that run missed or wrongly returned a match. The first three bars are starting points, and the five bars after them are the changes this article tests.

The chart shows where the money went. Starting from the $0.20 query starting point, the total drop was about $0.19. The pets field check accounts for 44 percent of it and the cheaper model with a backup for 39 percent. The shorter prompt accounts for 9 percent, and reused answers for 8 percent. Model time fell from 352 to 139 seconds.

Per 1,000 pairs, the estimated cost went from $0.080 to $0.0032. That is arithmetic on the rate card, and it does not forecast a bill for any other feed.

The first two steps down the chart, from sending everything to a city search and then to a full query, are ordinary query work. They account for most of the distance between $1.88 and $0.008, and they need no model at all.

What this test supports is narrow. It ranks the cost drivers for one agent on one sample of 2,500 pairs with 101 true matches and one run per configuration. It cannot say how the agent behaves for real renters, in other cities, on ads from another year, or on listings with longer text. It also cannot pick a universal best setting.

It supports one conclusion. In this agent most model calls were avoidable, and code could settle most of them. The cheaper model handled the rest at about 4 percent of the query starting point’s cost.

What should keep running after launch?

A recurring check should watch the same numbers I used here. Traffic, prompts, models, and retries all move cost after launch, and a trace makes each of them visible. Four numbers from this example are enough to start.

  • Model calls per pair. Sending every pair makes 1.0, the query starting point makes 0.106, and the final configuration makes 0.052. A jump means more pairs are reaching the model.

  • The share of pairs the field checks settle. It was 93.7 percent here, and a drop means the feed or the requirements changed.

  • Tokens per call, including reasoning tokens. A prompt change or a model change shows up here first.

  • Precision and recall on a fixed set of pairs with known answers. Rerunning these 2,500 pairs with the final configuration costs less than a cent at the rates above.

Two setup details will help. Register the rates with add_cost() so that Weave shows real numbers in its Usage table instead of zero. Keep the prompt version string in the cache key, so a prompt edit cannot reuse an old answer.

The confidence scores deserve a watch too. If almost every answer scores 5, an escalation rule that fires below 5 will rarely fire, and the backup model will sit idle while the cheaper model carries every decision.

Optimize for useful matches, not the smallest bill

The method in this article has four steps. Trace the run, find the model calls that code could settle, change one thing, and score the same pairs again. In this agent that sequence cut the estimated cost about 25 times against a database query starting point and left the answers intact. Against sending every pair, the cut was about 237 times, but few real agents start there. The same steps could have exposed a quality loss, and the runs were built to show one if it appeared.

Open one real trace from your own agent and count the calls that a field check or a lookup could have answered. Then test the largest cost driver first, on a fixed set of examples with answers you trust. The next technical question I would ask is whether a lower reasoning effort keeps the same answers at a lower price, since output tokens drove most of Luna’s bill.

Selected Sources

  1. FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance, Lingjiao Chen, Matei Zaharia, and James Zou, 2023. This paper names the three cost reduction strategies and reports the cascade result quoted in the opening. Its findings belong to the tasks the authors tested.

  2. Apartment for Rent Classified, UCI Machine Learning Repository, 2019, DOI 10.24432/C5X623, licensed under Creative Commons Attribution 4.0. This is the source of all 10,000 listings and the field descriptions.

  3. W&B Weave cost tracking documentation. It describes how Weave reads token usage, applies built in prices for supported integrations, and adds custom costs with add_cost().

  4. OpenAI API pricing, checked on September 27, 2026, for the Sol and Luna standard rates used in every cost estimate.

  5. OpenAI’s model pages for Sol and Luna, checked on September 27, 2026. Sol is described as built for complex coding and agentic workflows, and Luna as the most efficient model for focused, high volume tasks. The pages list gpt-6-sol as the default snapshot and gpt-6-luna as the single snapshot.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top