Datacorn

245 posts

Datacorn banner
Datacorn

Datacorn

@datacorn_io

Synthetic dataset generator. Supporting #developers to test software applications, #DataScience projects or #AI models.

For more, please visit 👉 Katılım Ağustos 2023
48 Takip Edilen9 Takipçiler
Miguel Ángel Durán
Miguel Ángel Durán@midudev·
¡Pedazo de biblioteca de Atlassian para Drag & Drop! ✓ Funciona en React, Vue, Angular y Svelte ✓ Usada por Trello, Jira y Confluence ✓ Tamaño pequeño: 4.7KB ✓ Compatible con móviles → npm install @atlaskit/pragmatic-drag-and-drop
Español
11
205
1.4K
41.8K
Alex Xu
Alex Xu@alexxubyte·
Almost every software engineer has used Linux before, but only a handful know how its Boot Process works :) Let's dive in.
Alex Xu tweet media
English
7
180
948
75.1K
Sahn Lam
Sahn Lam@sahnlam·
Load Balancer Basics Load balancers are essential components in modern application architectures, designed to distribute incoming traffic efficiently across multiple servers. Load balancers improve application performance, availability, and scalability. Traffic Distribution: Load balancers evenly distribute incoming traffic among a pool of servers, ensuring optimal resource utilization and preventing any single server from becoming overwhelmed. Algorithms like round-robin or least connections are used to select the most suitable server for each request. High Availability: If a server fails, the load balancer automatically redirects traffic to the remaining healthy servers. This ensures that the application remains accessible even in the event of server failures, minimizing downtime and improving overall availability. SSL Termination: Load balancers can handle SSL/TLS encryption and decryption, offloading this CPU-intensive task from backend servers. This improves server performance and simplifies SSL certificate management. Session Persistence: For applications that require maintaining user sessions on a specific server, load balancers support session persistence. They ensure that subsequent requests from a user are consistently routed to the same server, preserving session integrity. Scalability: Load balancers facilitate horizontal scaling by allowing easy addition of servers to the pool. As traffic increases, new servers can be provisioned, and the load balancer will automatically distribute the load across all servers, enabling seamless scalability. Health Monitoring: Load balancers continuously monitor server health and performance. They exclude unhealthy servers from the pool, ensuring that only healthy servers handle incoming requests. This proactive monitoring maintains optimal application performance. – Subscribe to our weekly newsletter to get a Free System Design PDF (158 pages): bit.ly/496keA7
Sahn Lam tweet media
English
2
111
446
26.8K
Tech Fusionist
Tech Fusionist@techyoutbe·
Q. Which Kubernetes object is responsible for managing load balancing and routing traffic to a set of pods? A. Deployment B. DaemonSet C. Ingress D. Service
English
5
1
8
1.6K
Tech Fusionist
Tech Fusionist@techyoutbe·
Q. Which command is used to import existing infrastructure into Terraform? A. terraform state import B. terraform apply C. terraform init D. terraform import
English
6
0
7
1.2K
LlamaIndex 🦙
LlamaIndex 🦙@llama_index·
Building an Advanced Research Agent on @databricks 📈 We gave a 90 minute workshop last week at the @Data_AI_Summit on building an advanced research assistant with two focus areas beyond the naive RAG setup. 1. Improving Data Quality: Ensuring good parsing, chunking, indexing modules that can handle complex data types. LlamaParse is well-suited for this. 2. Improving Query Complexity: Adding layers of agentic reasoning to progressively handle more sophisticated queries over your data. We’re excited to share the full set of slides and two workshop notebooks below. We use LLMs available through Databricks and local @huggingface embeddings - you can easily run this within your own Databricks environment. Full slides: docs.google.com/presentation/d… [Improving Data Quality] Building a financial RAG pipeline with LlamaParse: colab.research.google.com/drive/18RUkf8I… [Improving Query Complexity] Building an assistant over research papers: colab.research.google.com/drive/18RUkf8I…
LlamaIndex 🦙 tweet media
English
7
59
271
122K
Datacorn
Datacorn@datacorn_io·
@ezekiel_aleke Very interesting, thank you very much. We will take it into account
English
0
0
1
56
Ezekiel
Ezekiel@ezekiel_aleke·
Dear Data Analyst! Learn the LookUps NOW. Repost for others
Ezekiel tweet media
English
6
399
1.8K
338.9K
Datacorn
Datacorn@datacorn_io·
@freest_man Very interesting, thank you very much. We will take it into account
English
0
0
1
6
Sasi 📊📈
Sasi 📊📈@freest_man·
ETL vs. ELT: What's the difference? Let's understand with business Examples (This might be asked in your DS interview) Before that, you can understand what Extraction, Transformation and Loading are in this Tweet: x.com/freest_man/sta… ETL Extract, transform, and load (ETL) is a data pipeline used to collect data from various sources. It then transforms the data according to business rules and loads the data into a destination data store. The transformation work in ETL takes place in a specialized engine, and it often involves using staging tables to temporarily hold data as it is being transformed and ultimately loaded to its destination. ELT Extract, load, and transform (ELT) differs from ETL solely in where the transformation takes place. In the ELT pipeline, the transformation occurs in the target data store. Instead of using a separate transformation engine, the processing capabilities of the target data store are used to transform data. This simplifies the architecture by removing the transformation engine from the pipeline. Examples of ETL/ELT ETL: Your organization has started to explore more about historical trends and patterns of the data. Currently, the organization only has a transactional database (OLTP) for the product, and running these heavy analytical queries runs the risk of breaking the transactional database for the product. Thus, you decide to replicate the data on a database designed specifically for analytics (OLAP) to power these heavy queries while not risking the production transactions database. You build an ETL pipeline to replicate data from the transactional database to the analytical database, including some data transformations to make it easier to use. ELT: Your organization purchases third-party data from a vendor to supplement your organization's data. While this third-party data is useful, the vendor provides extremely messy tables (Which is usually the scenario). Since this data is for R&D purposes, there is no defined business logic available. Therefore, you deem it’s okay to dump this data into a data lake and give the data science team access to explore. You build an ELT pipeline that extracts the raw data from the third-party vendor and loads it into the data lake. After a few iterations, the data science team determines which parts of the data are valuable, and builds data transformations on top of the data in the data lake for their workflows.
Sasi 📊📈 tweet mediaSasi 📊📈 tweet mediaSasi 📊📈 tweet media
Sasi 📊📈@freest_man

🌟Understanding ETL/ELT: Manage Data wisely 🌟 ETL and ELT are two common methods used by organizations to handle their data efficiently. Both stand for: ✨ E - Extract This is the first step in both ETL and ELT. It involves fetching data from various sources, such as databases, spreadsheets, websites, or applications. Imagine collecting puzzle pieces from different places. ✨ T - Transform After the data is extracted, it often needs some cleaning and organizing. This step is called data transformation. Here, the data is structured and made consistent, just like fitting the puzzle pieces together, so they form a clear picture. ✨ L - Load Once the data is extracted and transformed, it is loaded into a central storage place called a data warehouse or a database. It's like putting the completed puzzle in a safe and easily accessible box. ✨Examples of ETL/ELT 👉 ETL: Your organization is starting to explore more about historical trends and patterns of your data. Currently, the organization only has a transactional database (OLTP) for the product, and running these heavy analytical queries runs the risk of breaking the transactional database for the product. Thus, you decide to replicate the data on a database designed specifically for analytics (OLAP) to power these heavy queries while not risking the production transactions database. You build an ETL pipeline to replicate data from the transactional database to the analytical database, including some data transformations to make it easier to use. 👉 ELT: Your organization purchases third-party data from a vendor to supplement your organization's data. While this third-party data is useful, the vendor provides extremely messy tables(Which is usually the scenario). Since this data is for R&D purposes, there is no defined business logic available. Therefore, you deem it’s okay to dump this data into a data lake and give the data science team access to explore. You build an ELT pipeline that extracts the raw data from the third-party vendor and loads it into the data lake. After a few iterations, the data science team determines which parts of the data are valuable, and build data transformations on top of the data in the data lake for their workflows. --- That's a wrap! Retweet if you liked this post and follow @feest_man for more Thanks!

English
3
62
245
19.8K
Datacorn
Datacorn@datacorn_io·
@Python_Dv Very interesting, thank you very much. We will take it into account
English
0
0
0
5
Datacorn
Datacorn@datacorn_io·
@DataScienceDojo Very interesting, thank you very much. We will take it into account
English
0
0
0
3
Datacorn
Datacorn@datacorn_io·
@learnk8s Very interesting, thank you very much. We will take it into account
English
0
0
0
7
LearnKube
LearnKube@learnk8s·
This article discusses the complexities of learning Kubernetes and suggests that understanding its API is the most straightforward approach The article also provides a guide on how to manipulate the API ➜ @talhakhalid101/simplest-way-to-learn-kubernetes-is-with-its-api-07e1fb8c6e0f" target="_blank" rel="nofollow noopener">medium.com/@talhakhalid10
LearnKube tweet media
English
1
8
47
4.3K
Datacorn
Datacorn@datacorn_io·
@ProfTomYeh Very interesting, thank you very much. We will take it into account
English
0
0
1
39
Tom Yeh
Tom Yeh@ProfTomYeh·
[Superposition] by Hand✍️ Superposition is a key property that differentiates a qubit in a quantum computer from a bit in a classical computer. A bit is always in one position. A qubit can be in many positions simultaneously, until it is measured. This “super” power to be in many positions at once is what enables certain quantum algorithms to achieve exponential speedup! How can we calculate superposition by hand? [1] Given ↳ 3 Qubits: 🟦 a, 🟧 b and 🟪 c. ↳ A quantum circuit with 5 operations involving these 3 qubits. ↳ Because each quantum operation creates two branches, we will show how this circuit is exploring 2 ^ 5 = 32 positions simultaneously. [2] 🟦 Set a ↳ Qubit a is simultaneously exploring two new branches: |0⟩ and |1⟩ ↳ a = |0⟩ means we want a to have a positive chance in |0⟩ and no chance in |1⟩. We write + in |0⟩ and o in |1⟩ to indicate this. ↳ Now, the quantum circuit is exploring 2 positions simultaneously. [3] 🟧 Set b ↳ Continuing each branch from the previous step, Qubit b is simultaneously exploring two new branches: |0⟩ and |1⟩. ↳ b = |1⟩ means we want b to have no chance in |0⟩ and a positive chance in |1⟩. We write o in |0⟩ and 1 in |1⟩ to indicate this. ↳ Now, the quantum circuit is exploring 2 x 2 = 4 positions simultaneously. [4] 🟦 H|a⟩ Hadamard Gate ↳ Continuing each branch from the previous step, Qubit a is simultaneously exploring two new branches: |0⟩ and |1⟩ by applying an Hadamard gate. ↳ The main purpose of the H gate is to create “superposition” using the following rules: + |0⟩ → + |0⟩ + |1⟩ + |1⟩ → + |0⟩ - |1⟩ ↳ Since a has + in |0⟩, the result is +’s for both |0⟩ and |1⟩. ↳ Intuitively, it is like split the + in the |0⟩ branch into two +’s, one for |0⟩ and one for |1⟩. ↳ Now, the quantum circuit is exploring 2 x 2 x 2 = 8 positions simultaneously. [5] 🟧 H|b⟩ Hadamard Gate ↳ Continuing each branch from the previous step, Qubit b is simultaneously exploring two new branches: |0⟩ and |1⟩ by applying an Hadamard gate. ↳ Recall the rules are: + |0⟩ → + |0⟩ + |1⟩ + |1⟩ → + |0⟩ - |1⟩ ↳ Since b has + in |1⟩, the result is + for |0⟩ and - for |1⟩. ↳ Now, the quantum circuit is exploring 2 x 2 x 2 x 2 = 16 positions simultaneously. [6] 🟪 Set c ↳ Qubit c is simultaneously exploring two new branches: |0⟩ and |1⟩ ↳ c = |0⟩ means we want c to have a positive chance in |0⟩ and no chance in |1⟩. We write + in |0⟩ and o in |1⟩ to indicate this. ↳ Now, the quantum circuit is exploring 2 x 2 x 2 x 2 x 2 = 32 positions simultaneously. [7] Possible Paths ↳ Even though this quantum circuit is exploring all 32 probable paths simultaneously, only some paths are possible. ↳ If a branch has a 0, it means it is not possible to take the path through the branch. ↳ In other words, if a path has + or - for all the branches along the way, it is possible. ↳ We found only four possible paths without any zero along the way: 000 010 100 110 [8] Quantum Measurement ↳ When taking a measurement, the quantum nature will randomly choose one of the possibilities. ↳ Since there are four possibilities, each possibility has an equal 25% chance of being measured. ↳ Once measured, they collapse into just one of these possibilities. Notes: Mathematically, the quantum circuit in this exercise can be calculated using the Dirac notation as follows: H|0⟩ ⊗ H|1⟩ ⊗ |0⟩ = (1/2) * ( |000⟩ + |010⟩ + |100⟩ + |110⟩) In developing this exercise, my goal is to invent a simpler and more accessible method to calculate quantum circuits by hand without resorting to difficult math such as Dirac notation, complex numbers, and tensor-product, while yielding the same results. I would appreciate your feedback. Your feedback will help me develop more exercises like this on other quantum computing topics such as entanglement, control gates, and quantum algorithms like Deutsch–Jozsa, Shor, Bernstein–Vazirani, and Glover.
English
23
154
768
56.2K
Datacorn
Datacorn@datacorn_io·
@milan_milanovic Very interesting, thank you very much. We will take it into account
English
0
0
1
74
Dr Milan Milanović
Dr Milan Milanović@milan_milanovic·
𝗬𝗼𝘂𝗿 𝗦𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗔𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲 𝗜𝘀 𝗔𝗹𝘄𝗮𝘆𝘀 𝗖𝗼𝗺𝗽𝗹𝗲𝘅 𝗔𝘀 𝗬𝗼𝘂𝗿 𝗢𝗿𝗴𝗮𝗻𝗶𝘇𝗮𝘁𝗶𝗼𝗻 Have you heard about 𝗖𝗼𝗻𝘄𝗮𝘆'𝘀 𝗟𝗮𝘄? It is a theory created by computer scientist Melvin Conway in 1967. which says: "𝘖𝘳𝘨𝘢𝘯𝘪𝘻𝘢𝘵𝘪𝘰𝘯𝘴, 𝘸𝘩𝘰 𝘥𝘦𝘴𝘪𝘨𝘯 𝘴𝘺𝘴𝘵𝘦𝘮𝘴, 𝘢𝘳𝘦 𝘤𝘰𝘯𝘴𝘵𝘳𝘢𝘪𝘯𝘦𝘥 𝘵𝘰 𝘱𝘳𝘰𝘥𝘶𝘤𝘦 𝘥𝘦𝘴𝘪𝘨𝘯𝘴 𝘸𝘩𝘪𝘤𝘩 𝘢𝘳𝘦 𝘤𝘰𝘱𝘪𝘦𝘴 𝘰𝘧 𝘵𝘩𝘦 𝘤𝘰𝘮𝘮𝘶𝘯𝘪𝘤𝘢𝘵𝘪𝘰𝘯 𝘴𝘵𝘳𝘶𝘤𝘵𝘶𝘳𝘦𝘴 𝘰𝘧 𝘵𝘩𝘦𝘴𝘦 𝘰𝘳𝘨𝘢𝘯𝘪𝘻𝘢𝘵𝘪𝘰𝘯𝘴." In other words, the 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 𝗼𝗳 𝗮 𝘀𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝘀𝘆𝘀𝘁𝗲𝗺 𝗶𝘀 𝗼𝗳𝘁𝗲𝗻 𝗶𝗻𝗳𝗹𝘂𝗲𝗻𝗰𝗲𝗱 𝗯𝘆 𝘁𝗵𝗲 𝘀𝘁𝗿𝘂𝗰𝘁𝘂𝗿𝗲 𝗮𝗻𝗱 𝗰𝗼𝗺𝗺𝘂𝗻𝗶𝗰𝗮𝘁𝗶𝗼𝗻 𝗽𝗮𝘁𝘁𝗲𝗿𝗻𝘀 𝘄𝗶𝘁𝗵𝗶𝗻 𝘁𝗵𝗲 𝘁𝗲𝗮𝗺 𝗯𝘂𝗶𝗹𝗱𝗶𝗻𝗴 𝗶𝘁. This can result in more optimal software architecture for the problem being solved, as the team may focus on their own organizational needs over the system's needs. This means 𝗮𝗻 𝗼𝗿𝗴𝗮𝗻𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝘄𝗶𝘁𝗵 𝘀𝗺𝗮𝗹𝗹 𝗱𝗶𝘀𝘁𝗿𝗶𝗯𝘂𝘁𝗲𝗱 𝘁𝗲𝗮𝗺𝘀 𝘄𝗶𝗹𝗹 𝗽𝗿𝗼𝗱𝘂𝗰𝗲 𝗮 𝗺𝗼𝗱𝘂𝗹𝗮𝗿 𝘀𝗲𝗿𝘃𝗶𝗰𝗲 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲, 𝘄𝗵𝗶𝗹𝗲 𝗮𝗻 𝗼𝗿𝗴𝗮𝗻𝗶𝘇𝗮𝘁𝗶𝗼𝗻 𝘄𝗶𝘁𝗵 𝗹𝗮𝗿𝗴𝗲 𝗰𝗼𝗹𝗹𝗼𝗰𝗮𝘁𝗲𝗱 𝘁𝗲𝗮𝗺𝘀 𝘄𝗶𝗹𝗹 𝗽𝗿𝗼𝗱𝘂𝗰𝗲 𝗮 𝗺𝗼𝗻𝗼𝗹𝗶𝘁𝗵𝗶𝗰 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲. In some broad sense, we could even say that 𝗛𝗥 𝘂𝘀𝘂𝗮𝗹𝗹𝘆 𝗱𝗲𝗳𝗶𝗻𝗲𝘀 𝘀𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗮𝗿𝗰𝗵𝗶𝘁𝗲𝗰𝘁𝘂𝗿𝗲𝘀. To mitigate this, we can use the 𝗜𝗻𝘃𝗲𝗿𝘀𝗲 𝗖𝗼𝗻𝘄𝗮𝘆 𝗺𝗮𝗻𝗲𝘂𝘃𝗲𝗿. This technique means we should involve software architects, engineers, and leaders in defining organizational structures. Doing it can lead to better software. Yet, we can see that many organizations ignore Conway's law and think that organizational structures and software architecture are detached from each other, with surprises in the end. Back to you, what are your experiences with Conway's Law? Image credits: Manu Cornet (bonkersworld .net). #softwarearchitecture
Dr Milan Milanović tweet media
English
10
101
601
89.5K
Datacorn
Datacorn@datacorn_io·
@parmardarshil07 Very interesting, thank you very much. We will take it into account
English
0
0
0
311
Darshil | Data Engineer👨🏻‍🔧
Becoming AWS Data Engineer (Quick Guide) 📄 There are 100s of services available on Amazon Web Services, the thing is you don't need to learn all of them. You need to focus on 10-20 services that you might use as a Data Engineer and services for other work (networking/access management) 📈 AWS Services for Data Engineers: ✅ Simple Storage Service (S3): It is an object storage (you can store anything) basically, the center of your work, all of the data you will get will be stored here for future processing. ✅ AWS Glue: Want to write ETL (Extract, Transform, Load) jobs in Python/Spark without worrying about servers? Glue is a serverless service where you just need to focus on writing your code and everything will be taken care of by AWS ✅ Amazon Redshift: So you processed data and wrote the ETL script, where to load it? Yes! Data Warehouse. Amazon Redshift is a fully managed data warehouse service that makes it simple to analyze all your data using standard SQL and your preferred business intelligence (BI) tools. ✅ Amazon EMR (Elastic MapReduce): Maybe you already have your processing scripts on on-premise servers written in Hadoop/Spark and want to migrate to a similar system then EMR is the way to go. It is a managed big data platform that makes it easy to process large amounts of data using open-source data processing frameworks like Apache Hadoop and Apache Spark. ✅ AWS Lambda: Want to run quick scripts on specific times/triggers/events? Lambda is your best friend! AWS Lambda is a serverless computing service that enables you to run your code without managing servers. This service can be used to run data engineering workflows and transform data in real-time. ✅ Amazon Athena: Why load data on Data Warehouse, when you can just direct run SQL query on top of actual files? Athena is an ad-hoc query interface that you can use for SQL queries directly on top of files stored on S3 buckets ✅ Kinesis: Want to process real-time data, analyze it and store it? Kinesis is used for this type of work, just like Apache Kafka, you can process real-time data using Kinesis. ✅ DMS (Data Migration Service): I spent my early data engineering days working on migration projects and DMS makes your life so much easier. There are many more services you can focus on such as EC2, IAM, VPC, Batch, Sagemaker, etc... If you want to learn all of these services by building a project then check out the comments 👇🏻
Darshil | Data Engineer👨🏻‍🔧 tweet media
English
5
92
570
55.8K
Datacorn retweetledi
Level Up Coding
Level Up Coding@LevelUpCoding_·
Database Indexing Explained Most databases require some form of indexing to keep up with performance benchmarks. Searching through a database is much simpler when the data is correctly indexed, which improves the system's overall performance. A database index is a lot like the index on the back of a book. It saves you time and energy by allowing you to easily find what you're looking for without having to flick through every page. Database indexes work the same way. An index is a key-value pair where the key is used to search for data instead of the corresponding indexed column(s), and the value is a pointer to the relevant row(s) in the table. To get the most out of your database, you should use the right index type for the job. The B-tree is one of the most commonly used indexing structures where keys are hierarchically sorted. When searching data, the tree is traversed down to the leaf node that contains the appropriate key and pointer to the relevant rows in the table. B-tree is most commonly used because of its efficiency in storing and searching through ordered data. Their balanced structure means that all keys can be accessed in the same number of steps, making performance consistent. Hash indexes are best used when you are searching for an exact value match. The key component of a hash index is the hash function. When searching for a specific value, the search value is passed through a hash function which returns a hash value. That hash value tells the database where the key and pointers are located in the hash table. Bitmap indexing is used for columns with few unique values. Each bitmap represents a unique value. A bitmap indicates the presence or absence of a value in a dataset, using 1’s & 0’s. For existing values, the position of the 1 in the bitmap shows the location of the row in the table. Bitmap indexes are very effective in handling complex queries where multiple columns are used. When you are indexing a table, make sure to carefully select the columns to be indexed based on the most frequently used columns in WHERE clauses. A composite index may be used when multiple columns are often used in a WHERE clause together. With a composite index, a combination of two or more columns are used to create a concatenated key. The keys are then stored based on the index strategy, such as the options mentioned above. Indexing can be a double-edged sword. It significantly speeds up queries, but it also takes up storage space and adds overhead to operations. Balancing performance & optimal storage is crucial to get the most out of your database without introducing inefficiencies.
Level Up Coding tweet media
English
15
190
822
88.9K