0% found this document useful (0 votes)

226 views

Apache Kafka Tutorial PDF

Uploaded by

Mahmoud Naser

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

0% found this document useful (0 votes)

226 views

Apache Kafka Tutorial PDF

Uploaded by

Mahmoud Naser

We take content rights seriously. If you suspect this is your content, claim it here.

Available Formats

Download as PDF, TXT or read online on Scribd

You are on page 1/ 13

Apache Kafka

About the Tutorial

Apache Kafka was originated at LinkedIn and later became an open sourced Apache project in
2011, then First-class Apache project in 2012. Kafka is written in Scala and Java. Apache Kafka
is publish-subscribe based fault tolerant messaging system. It is fast, scalable and distributed
by design.

This tutorial will explore the principles of Kafka, installation, operations and then it will walk you
through with the deployment of Kafka cluster. Finally, we will conclude with real-time
applications and integration with Big Data Technologies.

Audience
This tutorial has been prepared for professionals aspiring to make a career in Big Data Analytics
using Apache Kafka messaging system. It will give you enough understanding on how to use
Kafka clusters.

Prerequisites
Before proceeding with this tutorial, you must have a good understanding of Java, Scala,
Distributed messaging system, and Linux environment.

Copyright and Disclaimer

All the content and graphics published in this e-book are the property of Tutorials Point (I) Pvt.
Ltd. The user of this e-book is prohibited to reuse, retain, copy, distribute or republish any
contents or a part of contents of this e-book in any manner without written consent of the
publisher.

We strive to update the contents of our website and tutorials as timely and as precisely as
possible, however, the contents may contain inaccuracies or errors. Tutorials Point (I) Pvt. Ltd.
provides no guarantee regarding the accuracy, timeliness or completeness of our website or its
contents including this tutorial. If you discover any errors on our website or in this tutorial,
please notify us at contact@tutorialspoint.com

i
Apache Kafka

Table of Contents
About the Tutorial............................................................................................................................................... i

Audience ............................................................................................................................................................. i

Prerequisites ....................................................................................................................................................... i

Copyright and Disclaimer .................................................................................................................................... i

Table of Contents ............................................................................................................................................... ii

1. KAFKA – INTRODUCTION ............................................................................................................... 1

What is a Messaging System? ............................................................................................................................ 1

What is Kafka? ................................................................................................................................................... 2

2. KAFKA – FUNDAMENTALS.............................................................................................................. 4

3. KAFKA – CLUSTER ARCHITECTURE ................................................................................................. 7

4. KAFKA – WORKFLOW..................................................................................................................... 9

Workflow of Pub-Sub Messaging ....................................................................................................................... 9

Workflow of Queue Messaging / Consumer Group .......................................................................................... 10

Role of ZooKeeper............................................................................................................................................ 11

5. KAFKA – INSTALLATION STEPS ..................................................................................................... 12

Step 1: Verifying Java Installation .................................................................................................................... 12

Step 2: ZooKeeper Framework Installation ...................................................................................................... 13

Step 3: Apache Kafka Installation ..................................................................................................................... 15

Step 4: Stop the Server..................................................................................................................................... 16

6. KAFKA – BASIC OPERATIONS........................................................................................................ 17

Single Node-Single Broker Configuration ......................................................................................................... 17

List of Topics .................................................................................................................................................... 18

Single Node-Multiple Brokers Configuration .................................................................................................... 20

ii
Apache Kafka

Creating a Topic ............................................................................................................................................... 21

Basic Topic Operations ..................................................................................................................................... 22

Deleting a Topic ............................................................................................................................................... 23

7. KAFKA – SIMPLE PRODUCER EXAMPLE ........................................................................................ 24

KafkaProducer API ........................................................................................................................................... 24

Producer API .................................................................................................................................................... 25

Configuration Settings...................................................................................................................................... 25

ProducerRecord API ......................................................................................................................................... 26

SimpleProducer application ............................................................................................................................. 27

Simple Consumer Example ............................................................................................................................... 29

ConsumerRecord API ....................................................................................................................................... 30

ConsumerRecords API ...................................................................................................................................... 31

Configuration Settings...................................................................................................................................... 31

SimpleConsumer Application ........................................................................................................................... 32

8. KAFKA – CONSUMER GROUP EXAMPLE ....................................................................................... 34

9. KAFKA – INTEGRATION WITH STORM .......................................................................................... 37

About Storm .................................................................................................................................................... 37

Integration with Storm..................................................................................................................................... 37

Bolt Creation .................................................................................................................................................... 39

Submitting to Topology .................................................................................................................................... 42

Execution ......................................................................................................................................................... 44

10. KAFKA – INTEGRATION WITH SPARK............................................................................................ 45

About Spark ..................................................................................................................................................... 45

Integration with Spark ..................................................................................................................................... 45

11. KAFKA – REAL-TIME APPLICATION (TWITTER) .............................................................................. 50

iii
Apache Kafka

Twitter Streaming API ...................................................................................................................................... 50

12. KAFKA – TOOLS ............................................................................................................................ 55

System Tools .................................................................................................................................................... 55

Replication Tool ............................................................................................................................................... 55

13. KAFKA – APPLICATIONS ............................................................................................................... 56

iv
Apache Kafka
1. Kafka – Introduction

In Big Data, an enormous volume of data is used. Regarding data, we have two main
challenges. The first challenge is how to collect large volume of data and the second challenge
is to analyze the collected data. To overcome those challenges, you must need a messaging
system.

Kafka is designed for distributed high throughput systems. Kafka tends to work very well as
a replacement for a more traditional message broker. In comparison to other messaging
systems, Kafka has better throughput, built-in partitioning, replication and inherent fault-
tolerance, which makes it a good fit for large-scale message processing applications.

What is a Messaging System?

A Messaging System is responsible for transferring data from one application to another, so
the applications can focus on data, but not worry about how to share it. Distributed messaging
is based on the concept of reliable message queuing. Messages are queued asynchronously
between client applications and messaging system. Two types of messaging patterns are
available – one is point to point and the other is publish-subscribe (pub-sub) messaging
system. Most of the messaging patterns follow pub-sub.

Point to Point Messaging System

In a point-to-point system, messages are persisted in a queue. One or more consumers can
consume the messages in the queue, but a particular message can be consumed by a
maximum of one consumer only. Once a consumer reads a message in the queue, it
disappears from that queue. The typical example of this system is an Order Processing
System, where each order will be processed by one Order Processor, but Multiple Order
Processors can work as well at the same time. The following diagram depicts the structure.

5
Apache Kafka

Publish-Subscribe Messaging System

In the publish-subscribe system, messages are persisted in a topic. Unlike point-to-point
system, consumers can subscribe to one or more topic and consume all the messages in that
topic. In the Publish-Subscribe system, message producers are called publishers and message
consumers are called subscribers. A real-life example is Dish TV, which publishes different
channels like sports, movies, music, etc., and anyone can subscribe to their own set of
channels and get them whenever their subscribed channels are available.

What is Kafka?
Apache Kafka is a distributed publish-subscribe messaging system and a robust queue that
can handle a high volume of data and enables you to pass messages from one end-point to
another. Kafka is suitable for both offline and online message consumption. Kafka messages
are persisted on the disk and replicated within the cluster to prevent data loss. Kafka is built
on top of the ZooKeeper synchronization service. It integrates very well with Apache Storm
and Spark for real-time streaming data analysis.

Benefits
Following are a few benefits of Kafka:

 Reliability - Kafka is distributed, partitioned, replicated and fault tolerance.

 Scalability - Kafka messaging system scales easily without down time.

 Durability - Kafka uses “Distributed commit log” which means messages persists on
disk as fast as possible, hence it is durable.

 Performance - Kafka has high throughput for both publishing and subscribing
messages. It maintains stable performance even many TB of messages are stored.

6
Apache Kafka

Kafka is very fast and guarantees zero downtime and zero data loss.

Use Cases
Kafka can be used in many Use Cases. Some of them are listed below:

 Metrics - Kafka is often used for operational monitoring data. This involves
aggregating statistics from distributed applications to produce centralized feeds of
operational data.

 Log Aggregation Solution - Kafka can be used across an organization to collect logs
from multiple services and make them available in a standard format to multiple
consumers.

 Stream Processing - Popular frameworks such as Storm and Spark Streaming read
data from a topic, processes it, and write processed data to a new topic where it
becomes available for users and applications. Kafka’s strong durability is also very
useful in the context of stream processing.

Need for Kafka

Kafka is a unified platform for handling all the real-time data feeds. Kafka supports low latency
message delivery and gives guarantee for fault tolerance in the presence of machine failures.
It has the ability to handle a large number of diverse consumers. Kafka is very fast, performs
2 million writes/sec. Kafka persists all data to the disk, which essentially means that all the
writes go to the page cache of the OS (RAM). This makes it very efficient to transfer data
from page cache to a network socket.

7
Apache Kafka
2. Kafka – Fundamentals

Before moving deep into the Kafka, you must aware of the main terminologies such as topics,
brokers, producers and consumers. The following diagram illustrates the main terminologies
and the table describes the diagram components in detail.

In the above diagram, a topic is configured into three partitions. Partition 1 has two offset
factors 0 and 1. Partition 2 has four offset factors 0, 1, 2, and 3. Partition 3 has one offset
factor 0. The id of the replica is same as the id of the server that hosts it.

Assume, if the replication factor of the topic is set to 3, then Kafka will create 3 identical
replicas of each partition and place them in the cluster to make available for all its operations.
To balance a load in cluster, each broker stores one or more of those partitions. Multiple
producers and consumers can publish and retrieve messages at the same time.

8
Apache Kafka

Components Description

A stream of messages belonging to a particular category is called a

Topics
topic. Data is stored in topics.

Topics are split into partitions. For each topic, Kafka keeps a
minimum of one partition. Each such partition contains messages in
an immutable ordered sequence. A partition is implemented as a set
Partition of segment files of equal sizes.

Topics may have many partitions, so it can handle an arbitrary

amount of data.

Each partitioned message has a unique sequence id called as

Partition offset
“offset”.

Replicas are nothing but “backups” of a partition. Replicas are never

Replicas of partition
read or write data. They are used to prevent data loss.

i) Brokers are simple system responsible for maintaining the

published data. Each broker may have zero or more partitions per
topic. Assume, if there are N partitions in a topic and N number of
brokers, each broker will have one partition.

ii) Assume if there are N partitions in a topic and more than N brokers
(n + m), the first N broker will have one partition and the next M
Brokers broker will not have any partition for that particular topic.

iii) Assume if there are N partitions in a topic and less than N brokers
(n-m), each broker will have one or more partition sharing among
them. This scenario is not recommended due to unequal load
distribution among the broker.

Kafka’s having more than one broker are called as Kafka cluster. A
Kafka Cluster Kafka cluster can be expanded without downtime. These clusters are
used to manage the persistence and replication of message data.

9
Apache Kafka

Producers are the publisher of messages to one or more Kafka topics.

Producers send data to Kafka brokers. Every time a producer
publishes a message to a broker, the broker simply appends the
Producers
message to the last segment file. Actually, the message will be
appended to a partition. Producer can also send messages to a
partition of their choice.

Consumers read data from brokers. Consumers subscribes to one or

Consumers more topics and consume published messages by pulling data from
the brokers.

"Leader" is the node responsible for all reads and writes for the given
Leader
partition. Every partition has one server acting as a leader.

Node which follows leader instructions are called as follower. If the

leader fails, one of the follower will automatically become the new
Follower
leader. A follower acts as normal consumer, pulls messages and
updates its own data store.

10
Apache Kafka
3. Kafka – Cluster Architecture

Take a look at the following illustration. It shows the cluster diagram of Kafka.

11
Apache Kafka

End of ebook preview

If you liked what you saw…
Buy it from our store @ https://store.tutorialspoint.com

Heroku Cloud Application Development
From Everand
Heroku Cloud Application Development
Anubhav Hanjura
No ratings yet
Mastering Kafka Streams: From Basics to Expert Proficiency
From Everand
Mastering Kafka Streams: From Basics to Expert Proficiency
William Smith
No ratings yet
Apache Kafka Course Curriculum
No ratings yet
Apache Kafka Course Curriculum
5 pages
Getting Started With RabbitMQ and CloudAMQP
No ratings yet
Getting Started With RabbitMQ and CloudAMQP
137 pages
Complex Event Processing With Apache Flink Presentation
No ratings yet
Complex Event Processing With Apache Flink Presentation
49 pages
Pavel Chertorogov GraphQL The Holy Contract-Between Client and Server v11
No ratings yet
Pavel Chertorogov GraphQL The Holy Contract-Between Client and Server v11
95 pages
Apache Spark & Scala Course Content
No ratings yet
Apache Spark & Scala Course Content
5 pages
What's New in Angular 8?: Typescript 3.4 or Above Support. Ivy Renderer Engine Support
No ratings yet
What's New in Angular 8?: Typescript 3.4 or Above Support. Ivy Renderer Engine Support
7 pages
Analysis Node - Js Platform Web Application Security
No ratings yet
Analysis Node - Js Platform Web Application Security
60 pages
Full Reactive
No ratings yet
Full Reactive
53 pages
Getting Started With GKE
No ratings yet
Getting Started With GKE
44 pages
Spark SQL Tutorial
0% (1)
Spark SQL Tutorial
7 pages
Reactive Spring by Josh Long
100% (1)
Reactive Spring by Josh Long
378 pages
Understanding Unit and Integration Testing in Golang
No ratings yet
Understanding Unit and Integration Testing in Golang
59 pages
20 Best Practices For Working With Apache Kafka at Scale - DZone Big Data
No ratings yet
20 Best Practices For Working With Apache Kafka at Scale - DZone Big Data
10 pages
Apache Solr Search Patterns - Sample Chapter
No ratings yet
Apache Solr Search Patterns - Sample Chapter
33 pages
[FREE PDF sample] Python Unit Test Automation: Practical Techniques for Python Developers and Testers 1 / converted Edition Ashwin Pajankar ebooks
100% (2)
[FREE PDF sample] Python Unit Test Automation: Practical Techniques for Python Developers and Testers 1 / converted Edition Ashwin Pajankar ebooks
35 pages
GTS Enterprise Event Notification Rules Engine and Icon Selection
No ratings yet
GTS Enterprise Event Notification Rules Engine and Icon Selection
76 pages
Theangulartutorial PDF
No ratings yet
Theangulartutorial PDF
537 pages
Kafka Producer Internals: Find Answers On The Fly, or Master Something New. Subscribe Today
No ratings yet
Kafka Producer Internals: Find Answers On The Fly, or Master Something New. Subscribe Today
1 page
Databricks - Data Intelligence Platform For Advanced Data Architecture
No ratings yet
Databricks - Data Intelligence Platform For Advanced Data Architecture
5 pages
About Kubernetes and Security Practices - Short Edition: First Edition, #1
From Everand
About Kubernetes and Security Practices - Short Edition: First Edition, #1
Ami Adi
No ratings yet
Lecture Intro Kafka
No ratings yet
Lecture Intro Kafka
27 pages
Talend Open Studio For ESB Getting Started Guide
No ratings yet
Talend Open Studio For ESB Getting Started Guide
31 pages
Kubernetes For Beginners
100% (1)
Kubernetes For Beginners
29 pages
Feature Flag Best Practices
100% (1)
Feature Flag Best Practices
41 pages
API Design Done Right: Some Guidelines Which Every Programmer Should Probably Know
No ratings yet
API Design Done Right: Some Guidelines Which Every Programmer Should Probably Know
25 pages
Anusha Reddy
No ratings yet
Anusha Reddy
10 pages
Testing Microservices
No ratings yet
Testing Microservices
58 pages
TheInfoQeMag Service Mesh 1563784287455
No ratings yet
TheInfoQeMag Service Mesh 1563784287455
44 pages
Lambda Expressions With Collections Udemy
No ratings yet
Lambda Expressions With Collections Udemy
9 pages
Fastapi-Serviceutils: Release 2.0.0
No ratings yet
Fastapi-Serviceutils: Release 2.0.0
32 pages
Tekton Pipelines Master Course
No ratings yet
Tekton Pipelines Master Course
46 pages
Learning Azure DocumentDB
From Everand
Learning Azure DocumentDB
Becker Riccardo
No ratings yet
Behavior Driven Development BDDandContinuous Integration Delivery CI CD
No ratings yet
Behavior Driven Development BDDandContinuous Integration Delivery CI CD
13 pages
Mastering Concurrency Programming Java 8 Ebook B012o8s89k PDF
0% (1)
Mastering Concurrency Programming Java 8 Ebook B012o8s89k PDF
5 pages
Ultimate AWS Certified Solutions Architect Associate Exam Guide: Master Designing Resilient, Scalable Architectures with Core and Advanced AWS Services to Crack the SAA-C03 Certification (English Edition)
From Everand
Ultimate AWS Certified Solutions Architect Associate Exam Guide: Master Designing Resilient, Scalable Architectures with Core and Advanced AWS Services to Crack the SAA-C03 Certification (English Edition)
Venkata Sasi Kanumuri
No ratings yet
Public - Crash Course - Apache Spark - Berlin - 2018 PDF
No ratings yet
Public - Crash Course - Apache Spark - Berlin - 2018 PDF
76 pages
Kafka Internals
No ratings yet
Kafka Internals
30 pages
Building Scalable GraphQL APIs On AWS With CDK, TypeScript, AWS AppSync, Amazon DynamoDB, and AWS Lambda - AWS Mobile Blog
No ratings yet
Building Scalable GraphQL APIs On AWS With CDK, TypeScript, AWS AppSync, Amazon DynamoDB, and AWS Lambda - AWS Mobile Blog
17 pages
Kafka Cloudera Documentation
100% (1)
Kafka Cloudera Documentation
175 pages
Amazon Elastic Container Service
100% (1)
Amazon Elastic Container Service
683 pages
Lab - Exploring DataLake With Athena and Quicksight PDF
No ratings yet
Lab - Exploring DataLake With Athena and Quicksight PDF
22 pages
Getting Started With Knative
No ratings yet
Getting Started With Knative
81 pages
Large Scale Data Pipelines
No ratings yet
Large Scale Data Pipelines
91 pages
Overview of Deployment Options On AWS: June 2020
No ratings yet
Overview of Deployment Options On AWS: June 2020
21 pages
Sofware Engineering 82% Unified Modeling Language 80%
100% (1)
Sofware Engineering 82% Unified Modeling Language 80%
4 pages
From Monolith To Microservices
0% (1)
From Monolith To Microservices
2 pages
AWS API Gateway
100% (1)
AWS API Gateway
10 pages
12fa Docker Golang Sample
No ratings yet
12fa Docker Golang Sample
29 pages
Ebay Case Study
No ratings yet
Ebay Case Study
11 pages
Kubernetes A Complete Guide
From Everand
Kubernetes A Complete Guide
Gerardus Blokdyk
No ratings yet
Parallel Programming With Spark: Matei Zaharia
No ratings yet
Parallel Programming With Spark: Matei Zaharia
40 pages
How To Be A Good Software Architect
No ratings yet
How To Be A Good Software Architect
21 pages
Jenkins, Docker and Devops: The Innovation Catalysts: White Paper
No ratings yet
Jenkins, Docker and Devops: The Innovation Catalysts: White Paper
17 pages
Big Data Hadoop Training Certification 7
No ratings yet
Big Data Hadoop Training Certification 7
40 pages
Rule Engine
No ratings yet
Rule Engine
2 pages
Apache Maven Guide
No ratings yet
Apache Maven Guide
109 pages
Vue.js The Complete Reference: Mastering Modern Web Development with Vue 3, Composition API, and Scalable Patterns
From Everand
Vue.js The Complete Reference: Mastering Modern Web Development with Vue 3, Composition API, and Scalable Patterns
Aarav Joshi
No ratings yet
Apache Spark Theory by Arsh
No ratings yet
Apache Spark Theory by Arsh
4 pages
Cherrypy Tutorial PDF
0% (2)
Cherrypy Tutorial PDF
12 pages
Model Answer: Instructions To Students
No ratings yet
Model Answer: Instructions To Students
21 pages
Assessment 1 With Solution PDF
No ratings yet
Assessment 1 With Solution PDF
10 pages
Apache Derby Tutorial PDF
0% (1)
Apache Derby Tutorial PDF
15 pages
DHCP 23
No ratings yet
DHCP 23
6 pages
Generations of Computer
No ratings yet
Generations of Computer
7 pages
Question #1: Ans: B
No ratings yet
Question #1: Ans: B
25 pages
Less 05 SQL Trace and TKprof
No ratings yet
Less 05 SQL Trace and TKprof
16 pages
Logcat
No ratings yet
Logcat
634 pages
Rtos-Real-time Operating Systems: Echalone
No ratings yet
Rtos-Real-time Operating Systems: Echalone
39 pages
Risc in Pipe Ine
No ratings yet
Risc in Pipe Ine
39 pages
RCLSTG
No ratings yet
RCLSTG
7 pages
Atari em Fpga
No ratings yet
Atari em Fpga
53 pages
Ansible Automation Sibelius
No ratings yet
Ansible Automation Sibelius
3 pages
Study of The Dirty Copy On Write A Linux Kernel Memory Allocation Vulnerability
No ratings yet
Study of The Dirty Copy On Write A Linux Kernel Memory Allocation Vulnerability
6 pages
Microsoft Office
No ratings yet
Microsoft Office
1 page
Release (6.24.0000)
No ratings yet
Release (6.24.0000)
7 pages
Sony Optical Disc Archive Broschuere - E
No ratings yet
Sony Optical Disc Archive Broschuere - E
12 pages
Paper 1
No ratings yet
Paper 1
9 pages
GK Test-I
No ratings yet
GK Test-I
7 pages
Semaphore Twinsoft Manual
No ratings yet
Semaphore Twinsoft Manual
101 pages
HIKDDNS
No ratings yet
HIKDDNS
9 pages
TC-NC552S: 5MP Starlight Vandalproof Mini IR Dome Camera
No ratings yet
TC-NC552S: 5MP Starlight Vandalproof Mini IR Dome Camera
1 page
Computer Architecture - Memory System
100% (1)
Computer Architecture - Memory System
22 pages
IOT-SYLLABUS
No ratings yet
IOT-SYLLABUS
2 pages
Daftar Barang: Prioritas List Barang Kisaran Harga Jumlah Banyak Barang
No ratings yet
Daftar Barang: Prioritas List Barang Kisaran Harga Jumlah Banyak Barang
2 pages
Networking: Mike Pangburn
100% (1)
Networking: Mike Pangburn
27 pages
NodeManager SSLKeyException
No ratings yet
NodeManager SSLKeyException
2 pages
BASSYNTH Installation Instructions
No ratings yet
BASSYNTH Installation Instructions
5 pages
Intense PC2 Product Datasheet
No ratings yet
Intense PC2 Product Datasheet
3 pages
Introduction To Operating Systems PDF
No ratings yet
Introduction To Operating Systems PDF
67 pages
Acer Aspire 3935 (Wistron SM30) PDF
No ratings yet
Acer Aspire 3935 (Wistron SM30) PDF
45 pages
Atg 10 0 1p1
No ratings yet
Atg 10 0 1p1
1 page
Contents in Detail: Foreword by HD Moore Xiii Preface Xvii Acknowledgments Xix
No ratings yet
Contents in Detail: Foreword by HD Moore Xiii Preface Xvii Acknowledgments Xix
6 pages

Apache Kafka Tutorial PDF

Uploaded by

Apache Kafka Tutorial PDF

Uploaded by

Apache Kafka

About the Tutorial

Copyright and Disclaimer

Copyright and Disclaimer .................................................................................................................................... i

Table of Contents ............................................................................................................................................... ii

1. KAFKA – INTRODUCTION ............................................................................................................... 1

What is a Messaging System? ............................................................................................................................ 1

What is Kafka? ................................................................................................................................................... 2

3. KAFKA – CLUSTER ARCHITECTURE ................................................................................................. 7

Workflow of Pub-Sub Messaging ....................................................................................................................... 9

Workflow of Queue Messaging / Consumer Group .......................................................................................... 10

5. KAFKA – INSTALLATION STEPS ..................................................................................................... 12

Step 1: Verifying Java Installation .................................................................................................................... 12

Step 2: ZooKeeper Framework Installation ...................................................................................................... 13

Step 3: Apache Kafka Installation ..................................................................................................................... 15

Step 4: Stop the Server..................................................................................................................................... 16

6. KAFKA – BASIC OPERATIONS........................................................................................................ 17

Single Node-Single Broker Configuration ......................................................................................................... 17

List of Topics .................................................................................................................................................... 18

Single Node-Multiple Brokers Configuration .................................................................................................... 20

Creating a Topic ............................................................................................................................................... 21

Basic Topic Operations ..................................................................................................................................... 22

Deleting a Topic ............................................................................................................................................... 23

7. KAFKA – SIMPLE PRODUCER EXAMPLE ........................................................................................ 24

KafkaProducer API ........................................................................................................................................... 24

Producer API .................................................................................................................................................... 25

ProducerRecord API ......................................................................................................................................... 26

SimpleProducer application ............................................................................................................................. 27

Simple Consumer Example ............................................................................................................................... 29

ConsumerRecord API ....................................................................................................................................... 30

ConsumerRecords API ...................................................................................................................................... 31

SimpleConsumer Application ........................................................................................................................... 32

8. KAFKA – CONSUMER GROUP EXAMPLE ....................................................................................... 34

9. KAFKA – INTEGRATION WITH STORM .......................................................................................... 37

About Storm .................................................................................................................................................... 37

Integration with Storm..................................................................................................................................... 37

Bolt Creation .................................................................................................................................................... 39

Submitting to Topology .................................................................................................................................... 42

10. KAFKA – INTEGRATION WITH SPARK............................................................................................ 45

About Spark ..................................................................................................................................................... 45

Integration with Spark ..................................................................................................................................... 45

11. KAFKA – REAL-TIME APPLICATION (TWITTER) .............................................................................. 50

Twitter Streaming API ...................................................................................................................................... 50

12. KAFKA – TOOLS ............................................................................................................................ 55

System Tools .................................................................................................................................................... 55

Replication Tool ............................................................................................................................................... 55

13. KAFKA – APPLICATIONS ............................................................................................................... 56

What is a Messaging System?

Point to Point Messaging System

Publish-Subscribe Messaging System

 Reliability - Kafka is distributed, partitioned, replicated and fault tolerance.

 Scalability - Kafka messaging system scales easily without down time.

Need for Kafka

A stream of messages belonging to a particular category is called a

Topics may have many partitions, so it can handle an arbitrary

Each partitioned message has a unique sequence id called as

Replicas are nothing but “backups” of a partition. Replicas are never

i) Brokers are simple system responsible for maintaining the

Producers are the publisher of messages to one or more Kafka topics.

Consumers read data from brokers. Consumers subscribes to one or

Node which follows leader instructions are called as follower. If the

End of ebook preview

You might also like