---
title: "Responses API, structured extraction, and web search"
output: rmarkdown::html_vignette
vignette: >
  %\VignetteIndexEntry{Responses API, structured extraction, and web search}
  %\VignetteEngine{knitr::rmarkdown}
  %\VignetteEncoding{UTF-8}
---

```{r setup, include = FALSE}
fixture_dir <- "responses-api"
recording <- nzchar(Sys.getenv("FOUNDRY_RECORD_DOCS"))
have_fixtures <- dir.exists(fixture_dir) && length(list.files(fixture_dir)) > 0
run_api <- requireNamespace("httptest2", quietly = TRUE) &&
  (recording || have_fixtures)
library(foundryR)
if (run_api) {
  httptest2::start_vignette(fixture_dir)
}
knitr::opts_chunk$set(collapse = TRUE, comment = "#>", eval = run_api,
  fig.width = 7, fig.height = 4.5, out.width = "100%")
```

Calls to Azure show output recorded from a live run; setup code is shown but not run.

The Responses API is the foundryR route for stored response objects, stateful turns, built-in tools, token accounting, and schema-constrained output. The `model =` argument is a deployment name, and foundryR reads `AZURE_FOUNDRY_MODEL` when you omit it.

By default, responses go to the resource endpoint. Pass `project_endpoint =` to send one call to a Foundry project instead, or call `foundry_set_route("project")` to send responses, files, vector stores, and evaluations there for the rest of the R session. Agent-backed responses always use the project endpoint because agents live in a project.

```{r project-route, eval = FALSE}
foundry_set_project_endpoint(Sys.getenv("AZURE_FOUNDRY_PROJECT_ENDPOINT"))

foundry_response(
  "Summarize the project route in one sentence.",
  project_endpoint = Sys.getenv("AZURE_FOUNDRY_PROJECT_ENDPOINT")
)
```

## Response text and usage

`foundry_response()` returns a one-row tibble. Most analysis starts with the generated text and the token columns, not with the raw response.

```{r basic-response}
basic <- foundry_response(
  "Answer in one sentence: what is retrieval-augmented generation?"
)

basic$output_text
basic[, c(
  "input_tokens", "output_tokens", "reasoning_tokens",
  "cached_input_tokens", "total_tokens"
)]
```

Reasoning models can spend tokens that do not appear in `output_text`. Keep the token columns in reports when cost or model behavior matters.

## Stateful turns

Responses are stored by the service by default. Chaining with `previous_response_id` lets the service carry state from one turn to the next.

```{r stateful-turns}
first <- foundry_response(
  "Define catastrophic forgetting in one sentence."
)

second <- foundry_response(
  "Explain it for a college freshman in one sentence.",
  previous_response_id = first$response_id
)

second$output_text
second[, c("response_id", "input_tokens", "output_tokens", "total_tokens")]
```

Set `store = FALSE` for stateless calls when you do not need server-side state. Chaining requires a stored previous response.

## Structured extraction

Structured output is useful when model output becomes data. The example below codes short comments into sentiment, topic entities, and a summary. foundryR's schema helpers build the JSON Schema that Azure enforces.

```{r structured-extraction}
comment_schema <- foundry_schema(
  sentiment = schema_enum(c("positive", "negative", "neutral")),
  entities = schema_array(schema_string()),
  summary = schema_string()
)

comments <- c(
  "The new data pipeline reduced manual coding time by half.",
  "Participants reported confusion about the consent form."
)

comment_codes <- foundry_extract(
  comments,
  schema = comment_schema,
  schema_name = "CommentCode"
)

comment_codes[, c("sentiment", "entities", "summary", ".status")]
```

Top-level scalar fields become regular columns. Arrays and nested objects become list-columns, so you can unnest them only when your next analysis needs it.

If you already use ellmer type specifications, convert them locally with `as_foundry_schema()`, which returns the JSON Schema that foundryR sends with an extraction request. This keeps the extraction contract in one place.

```{r ellmer-schema, eval = requireNamespace("ellmer", quietly = TRUE)}
sentiment_spec <- ellmer::type_object(
  sentiment = ellmer::type_enum(
    c("positive", "negative", "neutral"),
    description = "Overall sentiment of the response."
  ),
  theme = ellmer::type_string("A short theme label for the response.")
)

sentiment_schema <- as_foundry_schema(sentiment_spec)
jsonlite::toJSON(sentiment_schema, auto_unbox = TRUE, pretty = TRUE)
```

## User-defined R tools

`foundry_tool()` describes an R function to the model and keeps the local function for execution. `foundry_agent()` runs a bounded loop: ask the model, execute requested function calls in R, send the matching tool outputs back, and stop when the model returns a final answer.

```{r function-tools}
get_weather <- function(location) {
  list(location = location, temperature = "70 F")
}

weather_tool <- foundry_tool(
  get_weather,
  description = "Get weather for a location",
  parameters = foundry_schema(
    location = schema_string("City and state.")
  )
)

tool_turns <- foundry_agent(
  "What is the weather in San Francisco?",
  tools = list(weather_tool),
  max_iterations = 4
)

tool_turns[, c("iteration", "final", "output_text")]
tool_turns$tool_calls[[1]][, c("type", "name", "call_id", "arguments")]
tool_turns$tool_results[[1]]
```

The maximum iteration count protects long jobs from unbounded tool loops. Set it to match the number of tool calls you are willing to review.

## Remote MCP tools

Microsoft documents remote Model Context Protocol tools for the Responses API. foundryR does not add a separate MCP helper because `foundry_response()` accepts raw Responses API tool objects.

```{r mcp-tool, eval = FALSE}
mcp_tool <- list(
  type = "mcp",
  server_label = "approved_server",
  server_url = Sys.getenv("MY_MCP_SERVER_URL"),
  require_approval = "never"
)

foundry_response(
  "Use the MCP server if it helps answer the question.",
  tools = list(mcp_tool)
)
```

Only attach MCP servers you trust and whose data handling your organization has approved. Treat the server as part of the same data boundary as the model call.

## Web-grounded answers

Web search sends query data to Grounding with Bing services. Microsoft documents that this can leave compliance or geographic boundaries and can incur extra cost, so avoid secrets and sensitive research data in web-search prompts. foundryR warns about this before the first web search in a session, and the option at the top of the next chunk acknowledges that warning.

`foundry_web_search()` requests the Responses API `web_search` tool and parses citations and tool calls into list-columns. The printed fields show the answer, sources, and search query separately.

```{r web-search}
options(foundryR.web_search_warning = TRUE)

web_answer <- foundry_web_search(
  "Which version of R does the R Project website list as the latest release, and when was it released?",
  search_context_size = "medium"
)

web_answer$output_text
web_answer$citations[[1]][, c("title", "url")]
web_answer$tool_calls[[1]][, c("type", "status", "action_type", "query")]
```

You can pass approximate location fields when the answer depends on place. Keep location values coarse unless the task needs more detail.

```{r web-search-location, eval = FALSE}
foundry_web_search(
  "Find a recent AI research event near me.",
  country = "US",
  region = "Washington",
  city = "Seattle",
  timezone = "America/Los_Angeles"
)
```

## Reasoning token accounting

`reasoning_effort` is sent as the Responses API `reasoning` object. The returned usage columns show whether hidden reasoning tokens contributed to cost.

```{r reasoning}
reasoned <- foundry_response(
  "Compare the two arguments and identify the weaker premise: A says the survey item is valid because it is short. B says it is valid because respondents interpret it consistently.",
  reasoning_effort = "medium"
)

reasoned$output_text
reasoned[, c(
  "input_tokens", "output_tokens", "reasoning_tokens",
  "cached_input_tokens", "total_tokens"
)]
```

## Streaming and chat completions

The Azure OpenAI Responses API supports Server-Sent Events streaming, but foundryR does not implement streaming. The package focuses on reproducible, tibble-returning analytical workflows. Use ellmer when you need interactive streaming chat in R.

Use `foundry_chat()` for the established chat-completions interface and simple assistant replies. Use `foundry_response()` for response IDs, built-in tools, structured output formats, richer output items, and token accounting.

```{r cleanup, include = FALSE}
if (run_api) {
  httptest2::end_vignette()
}
```
