---
title: "A Look at the DataFlow"
canonical: "https://onesaitplatform.refined.site/space/DOCT/2220820664/A%20Look%20at%20the%20DataFlow"
format: markdown
---
This module allows you to visually create and configure data flows between sources and destinations for both ETL/ELT-type processes and streaming flows, including among these flows transformations and data quality processes.

![image](media://1e0abf6c-b8b8-483d-8149-dbe2fd04db56)


Let's see a couple of examples:

- Ingest to Hadoop of the tail of a file with a process of field elimination:

![image](media://2314fc3b-a6f0-43a2-a4c8-40e53d211b4c)

- Ingest from a REST endpoint and load to the platform's Semantic DataHub with data quality process:

![image](media://bf418451-008c-4fa9-88f1-1bcafbd65675)


- It also offers a large number of connectors for specific communications, both input and output, as well as processors (in the Platform Developer Portal, you can see all the connectors: [http://bit.ly/2rwWZ1N](http://bit.ly/2rwWZ1N) ).

![image](media://1ff41151-1afc-4031-85fd-a976163fc246)

Among the main connectors of the Dataflow, you can find Big Data connectors with Hadoop, Spark, FTP, Files, Endpoint REST, JDBC, BD NoSQL, Kafka, Azure Cloud Services, AWS, Google, ...

- The platform-integrated component is the open-source software, StreamSets DataFlow ([https://streamsets.com](https://streamsets.com)) on which several connectors have been built to communicate with the platform:

![image](media://54a2f778-5fa8-4211-bfe3-e2ba9e631eab)

- All creation, development, deployment and monitoring of flows is performed from the platform's web console (ControlPanel):
- List of DataFlows by user with Administrator role (who can see flows from other users):

![image](media://b610d019-bab4-4647-8d13-bed810dc129f)

- DataFlow in development phase:

![image](media://cc1b936f-62a0-4c0e-81cc-14e4ded65b6d)

- DataFlow running:

![image](media://71299aed-2855-473f-94b7-3cb69b868cad)

- Debugging a DataFlow:

![image](media://84543142-0528-45d8-b44a-93e97df0959f)

- Fully integrated with the main Big Data technologies, including HDFS, HIVE, Spark, Kafka, SparkSQL, ... allowing to handle them easily and centrally:

![image](media://962a0880-6885-468d-be95-c5f58ca72062)

Besides connectors in areas such as IoT (OPC, CoAP, MQTT), Social Networks, ...

![image](media://22f46fe5-8259-4976-836e-2cf640c0b4cb)