> Markdown version of [/jobs/ext/70438-software-engineer-backend-python-content-understanding](https://www.wearedevelopers.com/jobs/ext/70438-software-engineer-backend-python-content-understanding). Every page supports `.md` or `Accept: text/markdown`. Links point to the HTML versions so they work for humans too. Agent guide: [/agents.md](https://www.wearedevelopers.com/agents.md). --- # Software Engineer (Backend, Python) - Content Understanding - **Company:** Scribd Inc. - **Location:** Portland, OR, United States (Remote available) - **Experience:** Experienced - **Salary:** $126,000.0 - $196,000.0 - **Contract:** Permanent contract - **Skills:** Application Programming Interfaces (APIs), Airflow, Amazon Web Services, Microsoft Azure, Program Optimization, Code Review, Content Analysis, Information Engineering, Distributed Systems, Hypertext Transfer Protocols (HTTP), Python (Programming Language), Machine Learning, Ruby on Rails, Ruby, Scala (Programming Language), Software Engineering, Datadog, Google Cloud, Amazon ElastiCache, Large Language Models, Apache Spark, AWS Lambda, Backend, Build Management, Infrastructure Automation Frameworks, Information Technology, Integration Frameworks, E-books, User Generated Content, Cloudwatch, Amazon Simple Queue Service (SQS), Terraform, Databricks - **Published:** May 14, 2026 - **Apply:** https://www.indeed.com/viewjob?jk=5c9a1ae81f793646 ## About the Role Do you have experience in System design?, Do you have a Bachelor's degree?, * 4+ years of professional software engineering experience * Proficiency in Python, Scala, Ruby, or similar languages * Experience designing and building distributed systems at scale * Hands-on experience building, deploying, and optimizing solutions using ECS, EKS, or AWS Lambda * Experience with infrastructure-as-code tools like Terraform (or similar) * Experience working with a public cloud provider (AWS, Azure, or Google Cloud) * Familiarity with data processing frameworks like Spark or Databricks for large-scale workloads * Proven ability to test, profile, and optimize systems for performance, scalability, and reliability * Bachelor's degree in Computer Science or equivalent professional experience * Bonus: Experience working with LLMs or integrating ML models into production systems ## Description We believe the best work happens when individual flexibility is balanced with meaningful community connection. Scribd Flex empowers employees to choose the workstyle and location that support their best performance, while committing to intentional in-person moments that strengthen collaboration and culture. Occasional in-person attendance is required for all Scribd, Inc. employees, regardless of location. So what are we looking for in new team members? At Scribd, Inc., we hire for "GRIT." Traditionally defined as the intersection of passion and perseverance toward long-term goals, GRIT reflects the mindset we expect from every employee. For us, it also serves as a practical framework for how we work: setting and achieving Goals, delivering Results within your role, contributing Innovative ideas and solutions, and strengthening the broader Team through collaboration and attitude. This posting reflects an approved, open position within the organization. About the team: The ML Content Understanding team powers metadata extraction, enrichment, and content understanding across all Scribd brands. We process hundreds of millions of documents, billions of images, and deliver high-quality metadata to enable content discovery and trust for millions of users worldwide. Our systems operate at massive scale, supporting diverse datasets like user-generated content (UGC), ebooks, audiobooks, and more. We work at the intersection of machine learning, data engineering, and distributed systems, collaborating closely with applied research and product teams to deploy scalable ML and LLM-powered solutions in production., We're seeking a Software Engineer II with strong backend development experience and a passion for solving complex data challenges at scale. In this role, you'll design, build, and optimize distributed systems that extract, enrich, and process metadata for a wide range of content. You'll work closely with ML engineers, product managers, and cross-functional partners to integrate machine learning models and LLM-based services into production pipelines and deliver impactful, high-performance solutions. This role offers the opportunity to work on cutting-edge generative AI and metadata enrichment problems at a truly global scale., Our team uses various technologies. The following are the ones that we use on a regular basis: Python, Scala, Ruby on Rails, Airflow, Databricks, Spark, HTTP APIs, AWS (Lambda, ECS, SQS, ElastiCache, Sagemaker, Cloudwatch, Datadog) and Terraform., * Design and build scalable systems to extract, enrich, and process metadata from millions of documents, images, and audio content. * Leverage LLMs to integrate capabilities like summarization, classification, extraction, and enrichment into metadata pipelines. * Collaborate with cross-functional teams, including ML engineers and product managers, to deliver scalable, efficient, and reliable metadata solutions. * Optimize and refactor existing systems for performance, scalability, and reliability. * Ensure data accuracy, integrity, and quality through automated validation and monitoring. * Participate in code reviews, ensuring best practices are followed and maintaining high-quality standards in the codebase. * Manage and maintain data pipelines, security and infrastructure ## Related Videos - [From Messy Queries to Scalable Systems - How Data Engineering actually works](https://www.wearedevelopers.com/videos/100203-from-messy-queries-to-scalable-systems-how-data-engineering-actually-works) - [Coffee with Developers - Maria Apazoglou](https://www.wearedevelopers.com/videos/1209-coffee-with-developers-maria-apazoglou) - [Coffee with Developers: David Heinemeier Hansson](https://www.wearedevelopers.com/videos/875-coffee-with-developers-david-heinemeier-hansson) - [Fireside Chat with Werner Vogels, VP & CTO, Amazon.com & Daniel Gebler, CTO at Picnic](https://www.wearedevelopers.com/videos/1405-fireside-chat-with-werner-vogels-vp-cto-amazon-com-daniel-gebler-cto-at-picnic) - [Shipping Faster with Less: Render on Cloud Hosting, AI Workloads, and the Future of DevOps](https://www.wearedevelopers.com/videos/1894-shipping-faster-with-less-render-on-cloud-hosting-ai-workloads-and-the-future-of-devops) - [AI Model Management Life Circles: ML Ops For Generative AI Models From Research to Deployment](https://www.wearedevelopers.com/videos/1152-ai-model-management-life-circles-ml-ops-for-generative-ai-models-from-research-to-deployment) ## Related Articles - [Highest Paying Tech Companies for Developers](https://www.wearedevelopers.com/magazine/220-highest-paying-tech-companies-for-developers) - [Dev Digest 121 - AI goes offline](https://www.wearedevelopers.com/magazine/456-dev-digest-121-ai-goes-offline) - [Dev Digest 120 - Apple and peers](https://www.wearedevelopers.com/magazine/455-dev-digest-120-apple-and-peers) - [Fully Remote Software Engineer Jobs](https://www.wearedevelopers.com/magazine/447-fully-remote-software-engineer-jobs) - [What Are Large Language Models?](https://www.wearedevelopers.com/magazine/304-what-are-large-language-models) - [Dev Digest 137 - AI'm not sure about this](https://www.wearedevelopers.com/magazine/485-dev-digest-137-ai-m-not-sure-about-this)