Principal-level software engineer with 20+ years of experience designing and building distributed systems, data platforms, and reliability-critical services. Ph.D. in Computer Science with deep expertise in performance analysis, scalable system design, and production reliability. Led systems processing 1B+ events per day at Oracle and built real-time infrastructure processing approximately 2 TB/hour at Amazon. Currently applying modern software engineering and AI-assisted development practices while re-architecting a mature research application in Go, React/TypeScript, PostgreSQL, and Docker. Seeking senior/staff backend or full-stack roles applying deep distributed-systems experience to complex, high-scale engineering problems.

gustavo.gomez@ggem.org -

CogLab Manager Modernization

Independent Software Engineer - Present

  • Re-architected and rewrote, using Go, React, TypeScript, PostgreSQL, and Docker, a university research lab-management application originally designed and built in 2005.
  • Designed a layered backend architecture separating HTTP handling, domain/business logic, and database access, using Go's net/http, chi, database/sql, and sqlc-generated type-safe queries.
  • Implemented domain logic for participant eligibility, experiment requirements, authentication, and constraint-based scheduling of research sessions against staff and resource availability.
  • Designed a PostgreSQL data model with foreign-key constraints, schema migrations, and row-level security as defense-in-depth for tenant isolation; deployed the application as self-hosted Docker services.
  • Used AI-assisted development as an engineering productivity tool, working in small increments and manually reviewing all generated code for correctness, maintainability, and architectural consistency.

Oracle

Site Reliability Developer -

  • Developed internal reliability tooling used by Oracle's Major Incident response organization to accelerate diagnosis, coordination, and mitigation of customer-impacting service outages.
  • Helped re-engineer the production Major Incident management platform using Oracle APEX, rapidly deploying the replacement alongside the legacy system before its full production cutover.
  • Participated in a 24/7 on-call rotation for high-severity incidents, coordinating technical response and serving as a certified Major Incident Communicator for executive and stakeholder communications.
  • Contributed to an experimental system applying large language models to operational logs to investigate event correlation and early indicators of Major Incidents.

Oracle

Principal Applications Engineer -

  • Lead engineer for distributed data services processing 1B+ daily events, optimizing for scalability, reliability, and performance.
  • Owned production and development Hadoop clusters, designing and operating large-scale MapReduce pipelines in Scala and data workflows using Hive and Bash.
  • Architected and delivered a major redesign of core services using Scala, Java, Kafka, Docker, Kubernetes, and Oracle OCI.
  • Designed fault-tolerant, high-availability architectures supporting mission-critical infrastructure.
  • Partnered cross-functionally to ensure end-to-end observability, reliability, and operational readiness of services.

University of Colorado. Psychology Department

Contract Software Developer (part-time) -

  • Built mobile data collection apps (iOS/Android) with real-time sync and offline caching.
  • Delivered cross-platform Ruby/Rails back-end services for large-scale behavioral studies.

Rosetta Stone

Software Developer -

  • Architected and implemented data collection and analytics pipelines for learner interaction data across millions of users.
  • Developed distributed RESTful APIs and media management systems handling large-scale concurrent access.
  • Built big-data integrations using Hadoop, HBase, Hive, and MapReduce.
  • Led CI/CD practices for multiple teams (Jenkins, Maven, Agile).

Amazon.com

Software Developer Engineer -

  • Designed and implemented real-time distributed systems for analyzing logs from thousands of servers (~2 TB/hour throughput).
  • Built tools to aggregate and suppress millions of system alarms across Amazon's compute fleet.
  • Contributed to low-latency, high-reliability systems foundational for large-scale ML and monitoring frameworks.
Boulder, CO
(303) 913-0725
gustavo.gomez@ggem.org
ggem.org

  • Languages: Go, Scala, Java, Python, C++, C, Ruby, TypeScript, JavaScript, Scheme
  • Backend & Data: REST APIs, PostgreSQL, MySQL, sqlc, relational data modeling
  • Frontend: React, TypeScript, JavaScript
  • Distributed Systems: Fault tolerance, scalable data platforms, high availability, concurrency
  • Cloud & Infrastructure: Docker, Kubernetes, Kafka, OCI, Hadoop, Hive, MapReduce, Linux
  • Reliability: Observability, incident response, performance analysis, profiling, optimization

  • Denver Museum of Nature and Science
    • Exhibit Guide
      -