> ## Documentation Index
> Fetch the complete documentation index at: https://docs.getcollate.io/llms.txt
> Use this file to discover all available pages before exploring further.

# REST API Hybrid Runner | Collate Integration Documentation

> Connect REST APIs to Collate using the Hybrid Runner. Deploy the ingestion agent in your environment for secure, private network metadata extraction.

export const MetadataIngestionUi = ({connector, selectServicePath, addNewServicePath, serviceConnectionPath}) => {
  return <>
   <p>
      To ingest metadata from your sources, you need to create a service connection.
      The service connects your source system with Collate. Once you create
      a service, you can use it to configure your ingestion workflows.<br />
      <br />
      To create a service connection and ingest your metadata, follow the steps below:
  </p>
  <Steps>
    <Step title="Select the Service">
    <ol>
          <li>
            On the left navigation bar, click <strong>Settings</strong>.
          </li>
          <li>
            On the next page, click <strong>Services</strong>, and then select the service.
            <img src="/public/images/connectors/visit-services-page.png" alt="Visit Services Page" />
          </li>
    </ol>
   </Step>


   <Step title="Create a New Service">

       To add a new service connection, click <strong>Add New Service</strong>.
      <img src="/public/images/connectors/create-new-service.png" alt="Create a new Service" />


   </Step>


     <Step title="Select the Connector">
       Select <strong>{connector}</strong> as the service type and click <strong>Next</strong>.


       {selectServicePath && <img src={selectServicePath} alt="Select Service" />}
   </Step>


   <Step title="Name and Describe your Service">
       Enter a unique, descriptive <strong>Service Name</strong> and <strong>Description</strong>.
       <ul>
         <li><strong>Service Name</strong>: Collate identifies services by their service name. Enter a name that distinguishes this deployment from other services, including other {connector} services you are ingesting metadata from.</li>
       </ul>


       <Note>
           The service name cannot be changed after it is set.
       </Note>


       {addNewServicePath && <img src={addNewServicePath} alt="Add New Service" />}
   </Step>


   <Step title="Configure the Service Connection">
       Set up the connection settings required for {connector}. <br /><br />

       Configure the following connection options to set up the service and start ingesting metadata from your sources. The right-hand panel displays help documentation for the selected connection type in the product UI

       {serviceConnectionPath && <img src={serviceConnectionPath} alt="Configure Service connection" />}
   </Step>
   </Steps>
   </>;
};

export const ConnectorDetailsHeader = ({name, icon, stage, availableFeatures, unavailableFeatures = [], availableFeaturesCollate = []}) => {
  const showSubHeading = availableFeatures?.length > 0 || unavailableFeatures?.length > 0 || availableFeaturesCollate?.length > 0;
  const totalAvailableFeatures = [...availableFeatures || [], ...availableFeaturesCollate || []];
  return <div className="container">
      <div className="Heading">
        <div className="flex items-center gap-3">
          {icon && <div className="IconContainer">
              <img src={icon} alt={name} noZoom className="ConnectorIcon" />
            </div>}
          <h1 className="ConnectorName">{name}</h1>
          <span className={`StageBadge ${stage === 'PROD' ? 'prod' : 'beta'}`}>
            {stage}
          </span>
        </div>
      </div>
      {showSubHeading && <div className="SubHeading">
          <div className="FeaturesHeading">Feature List</div>
          <div className="FeaturesList">
            {totalAvailableFeatures.map(feature => <div className="FeatureTag AvailableFeature" key={feature}>
                ✓ {feature}
              </div>)}
            {unavailableFeatures.map(feature => <div className="FeatureTag UnavailableFeature" key={feature}>
                ✕ {feature}
              </div>)}
          </div>
        </div>}
    </div>;
};

<ConnectorDetailsHeader icon="/public/images/connectors/rest.webp" name="REST" stage="PROD" availableFeatures={["API Endpoint", "Request Schema", "Response Schema"]} unavailableFeatures={[]} />

In this section, we provide guides and references to use the OpenAPI/REST connector.
Configure and schedule REST metadata workflows from the Collate UI:

* [Requirements](#requirements)
* [Metadata Ingestion](#metadata-ingestion)
* [Troubleshooting](/connectors/api/rest/troubleshooting)

## Requirements

Configure the schema source and file format before creating the REST service.

### Configure the OpenAPI Schema URL

1. Generate an [OpenAPI specification](https://swagger.io/specification/#openapi-document) for the service.
2. In the REST service connection form, select **OpenAPI Schema URL**.
3. In **OpenAPI Schema URL**, enter an HTTP or HTTPS URL that is reachable from the configured ingestion runner and points directly to the schema file.
4. **Optional:** In **Token**, enter a bearer token if the URL requires authentication.

#### Supported Formats

The connector supports these schema formats:

* JSON (`.json`)
* YAML (`.yaml` or `.yml`)

### Configure an S3-Hosted Schema

1. In the REST service connection form, select **OpenAPI Schema S3 URL** when the OpenAPI schema is stored in Amazon S3, including when the object is private.

2. In **OpenAPI Schema S3 URL**, enter the HTTPS object URL in the following format:

   ```text theme={null}
   https://<bucket>.s3.<region>.amazonaws.com/<object-key>
   ```

3. In **AWS Region**, enter the bucket's Amazon Web Services (AWS) Region.

4. In the AWS credential configuration, choose the credential method used by the ingestion runner:
   * Enable AWS Identity and Access Management (IAM) authentication with **IAM Auth** to use the runner's default AWS credential provider chain, including a workload or instance role.
   * Configure an access key and secret. Include the session token when using temporary credentials.
   * Select a named AWS profile that is available in the runner environment.

5. **Optional:** In `assumeRoleArn`, enter the target role Amazon Resource Name (ARN) after selecting a source credential method. The target role's trust policy must allow the source principal. For cross-account role assumption, the source principal's identity policy must also allow `sts:AssumeRole`. For same-account role assumption, an identity policy is required unless the trust policy grants the source principal permission directly.

Use the exact S3 object key in the URL. The connector recognizes `.json`, `.yaml`, and `.yml` extensions and detects the format of extensionless objects. Avoid percent-encoded characters in the object key because the connector passes the URL path to S3 without decoding it. The S3 source doesn't accept an `s3://` URI.

The REST connector retrieves the object with the configured AWS credentials. For more information about the URL structure, see [Virtual Hosting of General Purpose Buckets](https://docs.aws.amazon.com/AmazonS3/latest/userguide/VirtualHosting.html).

### Configure S3 Access for the Hybrid Runner

For a private schema object in the same AWS account as the effective ingestion principal, grant `s3:GetObject` through either the principal's identity policy or the bucket policy. A bucket policy isn't required when a same-account identity policy already grants access and no policy explicitly denies it.

For a cross-account schema object, configure both sides of the request:

1. Grant the effective ingestion principal `s3:GetObject` on the exact schema object through an AWS Identity and Access Management (IAM) identity policy.
2. Add a bucket policy that grants the same principal `s3:GetObject` on the exact schema object.

The following least-privilege bucket policy assumes that the effective principal is an IAM role. It identifies the role by its Amazon Resource Name (ARN). Replace the placeholders with the role and object used by ingestion:

```json theme={null}
{
  "Version": "2012-10-17",
  "Statement": [
    {
      "Sid": "AllowHybridRunnerOpenAPISchemaRead",
      "Effect": "Allow",
      "Principal": {
        "AWS": "arn:aws:iam::<account-id>:role/<effective-ingestion-role>"
      },
      "Action": "s3:GetObject",
      "Resource": "arn:aws:s3:::<bucket>/<object-key>"
    }
  ]
}
```

If static or profile credentials resolve to an IAM user, use the user's ARN as the bucket-policy principal and grant `s3:GetObject` to that user through an identity policy.

The `assumeRoleArn` setting selects the AWS credentials used by ingestion, but it doesn't grant S3 access. The target role becomes the effective ingestion principal, so it must have `s3:GetObject` permission. For cross-account S3 access, the bucket policy must also name the target role as the principal. For more information, see [Policies and Permissions in Amazon S3](https://docs.aws.amazon.com/AmazonS3/latest/userguide/access-policy-language-overview.html).

If the schema object uses server-side encryption with AWS Key Management Service (AWS KMS) keys, grant the effective principal `kms:Decrypt` permission on the key. Cross-account access requires a customer-managed KMS key because the AWS-managed `aws/s3` key can't be shared across accounts. The customer-managed key policy must also allow the effective principal.

This policy assumes that the bucket owner owns the schema object or uses **Bucket owner enforced** Object Ownership. If another account owns the object and access control lists (ACLs) are enabled, the object owner must grant read access or transfer ownership to the bucket owner.

## Metadata Ingestion

<MetadataIngestionUi connector={"REST"} selectServicePath={"/public/images/connectors/rest/select-service.png"} addNewServicePath={"/public/images/connectors/rest/add-new-service.png"} serviceConnectionPath={"/public/images/connectors/rest/service-connection.png"} />

### Connection Options

<Steps>
  <Step title="Configure the Schema Source">
    Configure one schema source:

    * **OpenAPI Schema URL**: Enter the HTTP or HTTPS location of the OpenAPI schema, such as `https://petstore3.swagger.io/api/v3/openapi.json`.
    * **OpenAPI Schema S3 URL**: Enter the HTTPS S3 object URL and configure the AWS Region and credentials described in [Configure an S3-Hosted Schema](#configure-an-s3-hosted-schema).

    **Optional:** In **Token**, enter a bearer token only when the **OpenAPI Schema URL** requires authentication.
  </Step>

  <Step title="Test the Connection">
    Once the credentials have been added, click on *Test Connection* and *Save* the changes.

    <img src="https://mintcdn.com/collatedocs/L7psA65ao88vmcRI/public/images/connectors/test-connection.png?fit=max&auto=format&n=L7psA65ao88vmcRI&q=85&s=2133f0d65f18df1e57f69d2cc3bdeff4" alt="Test Connection" width="1494" height="310" data-path="public/images/connectors/test-connection.png" />
  </Step>

  <Step title="Schedule the Ingestion and Deploy">
    Scheduling can be set up at an hourly, daily, weekly, or manual cadence. The
    timezone is in UTC. Select a Start Date to schedule for ingestion. It is
    optional to add an End Date.

    Review your configuration settings. If they match what you intended,
    click Deploy to create the service and schedule metadata ingestion.

    If something doesn't look right, click the Back button to return to the
    appropriate step and change the settings as needed.

    After configuring the workflow, you can click on Deploy to create the
    pipeline.

    <img src="https://mintcdn.com/collatedocs/piJyXg9wW6Ik1lg-/public/images/connectors/schedule.png?fit=max&auto=format&n=piJyXg9wW6Ik1lg-&q=85&s=f1add591824b44456f0e2ff259a21c6f" alt="Schedule the Workflow" width="2733" height="1083" data-path="public/images/connectors/schedule.png" />
  </Step>

  <Step title="View the Ingestion Pipeline">
    Once the workflow has been successfully deployed, you can view the
    Ingestion Pipeline running from the Service Page.

    <img src="https://mintcdn.com/collatedocs/cOe_QuHYxAbkMtTI/public/images/connectors/view-ingestion-pipeline.png?fit=max&auto=format&n=cOe_QuHYxAbkMtTI&q=85&s=8c754af74f99ee70e714f6f707b827e4" alt="View Ingestion Pipeline" width="2733" height="1271" data-path="public/images/connectors/view-ingestion-pipeline.png" />

    <Tip>
      If [AutoPilot](/how-to-guides/admin-guide/applications/autopilot) is enabled, workflows like usage tracking, data lineage, and similar tasks will be handled automatically. Users don’t need to set up or manage them - AutoPilot takes care of everything in the system.
    </Tip>
  </Step>
</Steps>

<Tip>
  When using a **Hybrid Ingestion Runner**, any sensitive credential fields—such as passwords, API keys, or private keys—must reference secrets using the following format:

  ```
  password: secret:/my/database/password
  ```

  This applies **only to fields marked as secrets** in the connection form (these typically mask input and show a visibility toggle icon).
  For a complete guide on managing secrets in hybrid setups, see the [Hybrid Ingestion Runner Secret Management Guide](https://docs.getcollate.io/getting-started/day-1/hybrid-saas/hybrid-ingestion-runner#3.-manage-secrets-securely).
</Tip>

## Troubleshooting

<Columns cols={2}>
  <Card title="REST Troubleshooting" href="/connectors/api/rest/troubleshooting">
    Learn more about how to troubleshoot common REST connector issues and resolve configuration or ingestion errors.
  </Card>
</Columns>
