Guide · Analytics

Run AWS Glue jobs locally

Waiting for a cloud job run to find a typo in a Glue script is a slow way to work. Goku, a free local AWS emulator, runs Glue code on your machine: ETL jobs and interactive notebooks on real Spark with the real awsglue libraries (GlueContext, DynamicFrame, Job), against a local Data Catalog and local S3.

Updated 27 September 2026 · Goku 4.0.0 · Free, no AWS account

Goku was formerly called Mimir. Same product, new name: the image is now tanujsoni027/goku (it was tanujsoni027/mimir-aws) and the command is goku. See old and new names.

On this page: Glue versions · Start · Run a job · Console and notebooks · Data Catalog and Athena · Libraries and formats · Differences · FAQ

Runtimes

Glue versions you can run locally

Pick the Glue version per job or per notebook, as in the AWS console. Every run happens in Docker on a real Glue runtime:

Glue versionSparkPythonJavaRuntime image
6.04.1.13.1317Built locally on first use
5.1 (default)3.5.63.1117AWS's published image
5.03.5.43.1117AWS's published image
4.03.3.03.108AWS's published image
3.0 / 2.03.1.1 / 2.4.33.78 / –AWS's published images

New jobs default to 5.1, which is what AWS uses for jobs created without a Glue version. The published images are pulled the first time you use a version, and they are large (5 to 7 GB each), so the first run takes a while. AWS publishes no Glue 6.0 image, so the first 6.0 job or notebook builds one from the official Glue 6.0 pieces. It downloads about 1 GB and takes roughly 3 minutes on a fast connection, and it only happens once.

Glue 6.0 behaves like Glue 6.0 on AWS: ANSI SQL mode is on, Python is 3.13, --extra-jars must be built for Scala 2.13, and getResolvedOptions rejects abbreviated argument names.

Step 1

Start Goku and turn on Glue

Install the goku CLI (see Install) and start Goku. The CLI mounts the Docker socket and the /tmp/goku-glue work folder that Glue runs need. Glue and Athena are not part of the core services, so subscribe to them once:

terminal
$ goku start
goku services subscribe glue athena
eval "$(goku env)"

You can also subscribe on the console's Services page, where the Data & analytics preset turns on Glue, Athena and related services together.

Step 2

Run a Glue job from the AWS CLI

Jobs in Goku are real Glue resources, so the AWS CLI, boto3, Terraform and Step Functions all see the same jobs and runs. Start with a small CSV file in local S3:

terminal
$ aws s3 mb s3://my-bucket
printf 'customer,amount\nalice,20\nbob,5\nalice,7\n' > orders.csv
aws s3 cp orders.csv s3://my-bucket/raw/orders/orders.csv

Then a Glue script that totals the amounts per customer and writes Parquet back to S3:

etl.py
import sys
from pyspark.context import SparkContext
from awsglue.context import GlueContext
from awsglue.job import Job
from awsglue.utils import getResolvedOptions

args = getResolvedOptions(sys.argv, ["JOB_NAME"])
glue_context = GlueContext(SparkContext.getOrCreate())
spark = glue_context.spark_session
job = Job(glue_context)
job.init(args["JOB_NAME"], args)

orders = spark.read.option("header", True).option("inferSchema", True).csv("s3://my-bucket/raw/orders/")
totals = orders.groupBy("customer").sum("amount").withColumnRenamed("sum(amount)", "total")
totals.show()
totals.write.mode("overwrite").parquet("s3://my-bucket/curated/totals/")

job.commit()

Upload the script, create the job and run it. The run log goes to CloudWatch Logs, as on AWS:

terminal
$ aws s3 cp etl.py s3://my-bucket/scripts/etl.py
aws glue create-job --name etl --role GlueRole --glue-version 5.1 \
  --command Name=glueetl,ScriptLocation=s3://my-bucket/scripts/etl.py \
  --worker-type G.1X --number-of-workers 2
RUN=$(aws glue start-job-run --job-name etl --query JobRunId --output text)
aws glue get-job-run --job-name etl --run-id $RUN --query JobRun.JobRunState
aws logs get-log-events --log-group-name /aws-glue/jobs/output --log-stream-name $RUN
aws s3 ls s3://my-bucket/curated/totals/

Run get-job-run again until the state is SUCCEEDED. The run goes through STARTING (where the image pull or the 6.0 build shows), then RUNNING. Errors land in /aws-glue/jobs/error. To try Glue 6.0, create the job with --glue-version 6.0.

Step 3

Or build it in the console, with notebooks

Open Glue in the console at http://localhost:8080. Under ETL jobs, Create job gives you a script editor, the Glue version picker, job parameters and run history with live logs. Under Notebooks, a new notebook starts an interactive session with a live, stateful kernel where spark and sc are already defined, which is the fastest way to explore data before writing a job.

AWS Glue ETL jobs in the Goku console: Spark and Python shell jobs running locally on the official Glue runtime image
Glue in the Goku console: jobs, notebooks, the Data Catalog, and crawlers, triggers and workflows.
Step 4

The Data Catalog, crawlers and Athena

Goku emulates the Glue Data Catalog: databases, tables and partitions. Create them with the AWS CLI, or point a crawler at a prefix in local S3. Crawlers classify CSV, JSON, Parquet, Avro and ORC files and turn key=value folders into partitions.

terminal
$ aws glue create-database --database-input '{"Name": "analytics"}'
aws glue create-table \
  --database-name analytics \
  --table-input '{
    "Name": "totals",
    "StorageDescriptor": {
      "Location": "s3://my-bucket/curated/totals/",
      "InputFormat": "org.apache.hadoop.hive.ql.io.parquet.MapredParquetInputFormat",
      "SerdeInfo": {
        "SerializationLibrary": "org.apache.hadoop.hive.ql.io.parquet.serde.ParquetHiveSerDe"
      },
      "Columns": [
        {"Name": "customer", "Type": "string"},
        {"Name": "total",    "Type": "bigint"}
      ]
    }
  }'

With the job parameter --enable-glue-datacatalog (the console sets it for new Spark jobs, and notebooks use it by default), spark.sql("SELECT ... FROM analytics.totals"), CREATE TABLE, INSERT INTO and glue_context.create_dynamic_frame.from_catalog read and write the local catalog. Athena uses the same catalog, so you can query a job's output with SQL in the console's query editor.

Amazon Athena query editor in the Goku console running SQL against Glue Data Catalog tables
Athena in Goku runs SQL over your Glue catalog tables and the S3 data behind them.
Real-world jobs

Libraries, data lake formats and orchestration

Know the limits

Differences from AWS Glue

Goku is a local sandbox, so a few things behave differently:

FAQ

Frequently asked questions

Can I run AWS Glue jobs locally without an AWS account?

Yes. Goku runs Glue jobs and notebooks in Docker on your machine against a local Data Catalog and local S3. It needs Docker, not an AWS account.

Which AWS Glue versions can I run locally?

Glue 6.0, 5.1 (the default), 5.0, 4.0, 3.0 and 2.0, chosen per job or notebook. 5.1 and older use AWS's published runtime images; 6.0 is built locally the first time you use it.

Why is the first Glue run slow?

The first run of a version pulls its runtime image, which is 5 to 7 GB, or for Glue 6.0 builds one once (about 1 GB to download, roughly 3 minutes on a fast connection). Later runs start straight away.

Can I use GlueContext, DynamicFrame and from_catalog locally?

Yes. Jobs run with the real awsglue libraries from the Glue runtime, and from_catalog reads tables from Goku's local Data Catalog.

Do Glue interactive notebooks work locally?

Yes. Notebooks in the console start an interactive session with a live Spark or Python kernel on the Glue version you pick.