Company Overview: Join our team at Accrete (https://www.accrete.ai), a product-focused AI company building intelligence-driven platforms that help organizations understand, assess, and act on complex information at scale using GenAI and agentic approaches. We focus on solving
• Building efficient storage for structured and unstructured data • Transform and aggregate the data using data processor technologies • Developing and deploying distributed computing Big Data applications using Open Source frameworks like Apache Spark, Apex, Flink,
• Experience with Hadoop and the HDFS Ecosystem • Strong Experience with Apache Spark, Storm, Kafka is a must. • Experience with Python, R, Pig, Hive, Kafka, Knox, Tomcat and Ambari • Experience with MongoDB • A
• Azure Data Factory • Azure Databricks • Python, Scala, PySpark, Spark • HIVE / HIVE LLAP / HBASE / CosmoDb • Azure Active Directory Domain Services • Apache Ranger / Apache Ambari • Azure Key Vault • Expertise
Candidate should be able to: Coordinate Development, Integration, and Production deployments. Optimize Spark code, Impala queries, and Hive partitioning strategy for better scalability, reliability, and performance. Build applications using Maven, SBT and integrated with continuous integration